On September 2, 2026, @OfficialLoganK announced Gemini 3.8 Flash as a new update focused on agentic and coding capabilities. The post describes it as the third updated Flash model in six weeks, while accompanying graphics compare its reported performance and token pricing with several other models.

Original post on X

Load the post to view it as published on X. X may receive connection data.

View the original post on X ↗

Gemini 3.8 Flash focuses on agentic work

The announcement is brief, but its emphasis is clear: Gemini 3.8 Flash is being positioned for tasks in which a model must do more than generate a single response. The post specifically calls out agentic and coding capabilities, suggesting a focus on workflows that involve multiple steps, tool use, software engineering, or interaction with external environments.

That positioning is reflected in the attached benchmark material. The graphics cover software engineering, terminal-based coding, computer use, financial analysis, legal workflows, document comprehension, chart reasoning, video understanding, and multidisciplinary reasoning. These categories are more representative of applied workflows than a single general-purpose score, although each result still reflects the particular benchmark and testing setup shown in the material.

Reported benchmark results span coding and professional tasks

The comparison graphic reports an 89.4% result for Gemini 3.8 Flash on Terminal-bench 2.1, labelled “agentic terminal coding.” Gemini 3.7 Flash is shown at 85.8%, while the other compared models range from 80.4% to 89.1%.

On Val’s Finance Agent v2, labelled “financial analyst tasks,” Gemini 3.8 Flash is shown at 61.4%. The graphic lists Gemini 3.7 Flash at 59.0% and the other models between 53.8% and 58.6%.

The attached results also show Gemini 3.8 Flash at 10.0% on Harvey’s Legal Agent Benchmark for complex legal workflows, compared with 8.8% for Gemini 3.7 Flash. On HLE-Verified, described as multidisciplinary expert reasoning, it is listed at 54.9%, narrowly ahead of Gemini 3.7 Flash at 53.6%.

Other reported figures include 86.2% on CharXiv Reasoning for information synthesis from complex charts and 86.2% on LABBench2 for biology real-world research tasks. On LVBench, the graphic shows separate agentic and static results for Gemini 3.8 Flash: 87.8% and 87.1%, respectively. These figures should be read as benchmark-specific results rather than a universal ranking across every type of AI task.

Introductory pricing is listed below regular rates

The pricing table in the accompanying material lists Gemini 3.8 Flash at $0.75 per million input tokens and $3.75 per million output tokens without caching. It shows regular prices of $1.50 per million input tokens and $7.50 per million output tokens in parentheses.

The same table states that the introductory pricing for Gemini 3.7 Flash and Gemini 3.8 Flash expires on December 31, 2026. It lists the regular rates as applying from January 1, 2027. That makes cost an important part of the model’s positioning for developers evaluating agentic workloads, where repeated model calls can make token economics consequential.

The chart cites a DeepMind evaluation-methodology page for the Gemini 3.8 Flash results. The supplied announcement does not provide additional details about how the model was trained, its deployment availability, or the precise conditions behind every comparison, so the benchmark figures are best understood as the results presented in the launch material.

What the announcement signals for developers

Gemini 3.8 Flash’s launch message combines three selling points: a rapid update cycle, an emphasis on agentic and coding tasks, and introductory token pricing. The attached comparisons suggest that Google is presenting the Flash line as a practical option for workflows that mix software development with structured professional or research tasks.

For developers, the most relevant question will be how those reported results translate to their own workloads. The announcement provides a concise performance snapshot, but selecting a model still depends on the specific tools, prompts, latency requirements, workload size, and reliability expectations involved.