Gemini 2.5 Flash

Fast Gemini model for low-latency multimodal and high-throughput tasks

Published
scheduleReleasedMarch 20, 2025

Gemini 2.5 Flash is optimized for speed and efficiency, making it suitable for interactive products, lightweight reasoning and high-volume calls.

starsCapabilities

visibilityVision understandingcodeFunction callingdata_objectStructured output

paymentsContext and pricing

Context limit1,048,576
Max output65,536
Knowledge cutoff2025-01
Input price$0.3/ 1M tokens
Output price$2.5/ 1M tokens
Cached input price$0.03/ 1M tokens

descriptionOverview

Overview

Gemini 2.5 Flash balances Gemini capabilities with faster response and better cost efficiency.

Best for

Use it for chat products, document helpers, extraction and real-time interactions.

lightbulbUse cases

  • Fast chat experiences
  • Lightweight document analysis
  • Extraction workflows
  • Real-time interactions

thumb_upStrengths

  • Low latency
  • Good cost profile
  • Multimodal support
  • Useful for high throughput

infoLimitations

  • Less capable than Pro models
  • May struggle with hardest reasoning tasks
  • Provider behavior can evolve quickly

compare_arrowsAlternative models

linkReferences

This content is compiled from official documentation and public sources. Always refer to official documentation for final details