Gemini 3.8 Flash Takes the Price War Down to Token Consumption

Modelos/proveedores relacionados: Claude Anthropic Gemini Google DeepMind GPT OpenAI Anthropic Proveedor Google DeepMind Proveedor OpenAI Proveedor
Gemini 3.8 Flash Takes the Price War Down to Token Consumption

Google released Gemini 3.8 Flash on September 2, 2026, keeping its promotional price at $0.75 per million input tokens and $3.75 per million output tokens. However, Artificial Analysis found that completing the same task cost about 40% more because the model consumed roughly 30% more output tokens.

Article image

The unchanged price table therefore does not necessarily mean an unchanged bill.

Article image

On DeepSWE, Gemini 3.8 Flash scored 73.7%, just 0.3 points below Claude Opus 5’s 74.0%, while costing less than one-sixth as much. Its trade-off is higher token consumption per task, with the promotional price scheduled to double at the start of next year. As agentic reasoning becomes deeper, the relevant measure is shifting from price per token to cost per completed task.

Specifications and pricing

Gemini 3.8 Flash is Google’s third Flash release in 43 days: version 3.6 arrived in late July, 3.7 on August 13, and 3.8 on September 2. CEO Sundar Pichai has said the series is intended to approach a monthly release cycle.

Its highlighted benchmarks cover four areas: 73.7% on DeepSWE v1.1 long-horizon software engineering; 54.9% on HLE-Verified, ahead of Opus 5 at 54.4% and GPT-5.6 Sol at 54.5%; 61.4% on the Vals financial-agent benchmark, ranking first; and approximately 305 tokens per second, reportedly the fastest independently measured commercial model.

Article image
Article image

The promotional price of $0.75 per million input tokens and $3.75 per million output tokens is valid through December 31. Comparable prices are $5/$25 for Opus 5 and $4/$20 for GPT-5.6 Sol. Gemini 3.8 Flash is available in the Gemini App, AI Studio, and the Antigravity coding platform. Cursor and Vercel integrated it on launch day.

Why token usage increased

Google describes the upgrade as the model working harder to complete tasks. In practice, this means deeper agentic reasoning: more reasoning steps and more frequent tool calls and iterations for complex jobs.

Article image

The model can test each step, inspect outputs, and revise its work, reducing final errors. This repeated reasoning, tool-use, and verification cycle consumes additional output tokens. Artificial Analysis measured about 30% more output tokens per task, translating into roughly 40% higher total task cost.

Even after the increase, the estimated cost of about $0.58 per task remains near the frontier of the intelligence-to-cost trade-off. Its high-reasoning mode reached a composite intelligence score of 59, three points above the previous generation.

Article image

After the promotional period, prices are expected to rise to $1.50 per million input tokens and $7.50 per million output tokens. Google will continue offering Gemini 3.7 Flash as a lower-cost option.

Competition for developer workloads

The rapid Flash release cycle reflects a contest for developer entry points. Coding agents and automated workflows are among the fastest-growing sources of token consumption and a major API revenue battleground. Because migration costs are low, buyers can switch based on intelligence delivered per unit cost.

Gemini 3.8 Flash’s near-parity with Opus 5 on DeepSWE, at less than one-sixth of the price, sharply lowers the market price for near-flagship coding capability.

Google also released Gemini 3.8 Flash Cyber. It scored 47.2% on CWE-Bench vulnerability repair, broadly matching leader Fable 5. Chrome’s security team reported 2.6 times more effective patches than larger commercial models, while Wiz’s internal penetration tests showed recall 7.5 to 9.7 percentage points higher at one-half-point-three to one-fifth-point-two of competitors’ cost.

Article image

The cybersecurity version is not publicly sold. It is available only to about 650 trusted organizations through the Fairwind program, a gated-access approach similar to Anthropic’s Mythos 5.1 trusted-access plan. Advanced cybersecurity capabilities are increasingly being treated as resources requiring qualification.

The apparent strategy is to use Flash for scale, developer adoption, and platform lock-in, while reserving higher-margin Pro products for a point when the capability gap is wider. This may also explain why the Gemini 3.5 Pro announced in June has not yet launched: a stronger Flash model makes it difficult to establish a viable Pro price anchor.

Where the model falls short

Google’s four highlighted benchmarks all show leadership, but third-party checks of the broader results are less consistent. On open-ended agent tasks, Opus 5 scored 51.8% on Terminal-Bench 4.0 versus 19.1% for Gemini 3.8 Flash. On OSWorld computer operation, the scores were 75.4% and 59.0%, respectively.

On the Harvey legal benchmark, the reported 10.0% lead is a relative result in which all models performed below a passing level. The general pattern is that Gemini 3.8 Flash can approach flagship performance on specialized tasks with clear boundaries and defined endpoints, but falls further behind on open tasks requiring autonomous planning. Wharton professor Ethan Mollick’s preliminary assessment was that it is a very good Flash model, but not a frontier model.

Article image
Article image

All comparison figures use vendor testing methodologies, and launch-day benchmarks often differ from production performance. The cost impact also depends on task structure: short question-and-answer interactions are barely affected, while long agentic workflows bear most of the increase. Reports that some teams are considering returning to Gemini 3.7 Flash to control budgets remain unverified.

Implications

The confirmed product facts are a Flash release cycle compressed to roughly three weeks, capabilities approaching flagship levels, and a cybersecurity edition controlled through eligibility requirements.

More broadly, AI price competition is moving beyond published token rates. When providers’ per-token prices converge, the number of tokens required to finish a task becomes a new pricing lever. Deeper reasoning and more steps make a model more expensive even when its price list remains unchanged. Anthropic’s 75% cut to its cache-read price two days earlier reflects the same broader competition over actual usage costs.

For developers, the practical lesson is to compare cost per completed task rather than token prices alone. The next phase of the price war may be determined by whose bills conceal the most consumption.

Compartir este artículo