In a post on X, @JustLingonberry presents an alleged Gemini 4 Pro leak as a major competitive threat to Anthropic and OpenAI models. The accompanying comparison image shows Gemini 4 Pro with higher displayed figures than Claude Opus 5.5 and GPT-6 Astra on each listed benchmark, along with lower displayed token prices. However, the material does not include an official Google announcement, model documentation, testing methodology or pricing page.

The thread is therefore best read as a report of what one comparison image appears to show—not as confirmation that Gemini 4 Pro has launched, achieved the listed results or will be sold at those prices.

Comparison image lists reported benchmark, context-window and displayed pricing figures for three models.
Comparison image lists reported benchmark, context-window and displayed pricing figures for three models.

Image credit: @JustLingonberry on X

What the X thread is reporting

The opening post says that Gemini 4 Pro “leaks” dominate both Anthropic and OpenAI and suggests that competing models may need to match its pricing. Replies add the author’s view that the leak appears believable and that the model may match or exceed Claude Opus 5.5.

Those statements are interpretations from the thread, not results from a disclosed evaluation. The thread also speculates that Google’s control over its own Tensor Processing Units (TPUs)—specialized chips for running AI workloads—could help with speed and pricing, but it supplies no evidence showing that the chips produced an advantage. There are also no measurements of inference speed, operating costs or margins.

The opening post refers to “Fable 5.5” and “GPT 6.1,” while the comparison image labels the competitors Claude Opus 5.5 and GPT-6 Astra. The source does not establish whether these are alternate names for the same models or different references.

Reported benchmark comparison

The attached image lists the following figures:

Benchmark

Gemini 4 Pro (leak)

Claude Opus 5.5

GPT-6 Astra

DeepSWE v1.1

88.7

74.2

74.1

Terminal-Bench 2.1

95.3

~88

~88

OSWorld 2.0

86.8

81.8 (partial)

72.6 (offline partial)

GDPval-AA

2,064 (v2)

1,846 (v2.1)

1,542 (v2.1)

The image displays a higher Gemini figure in every listed row. That describes the comparison graphic, not a verified model lead: the rows use different benchmark versions and include qualifications such as “partial” and “offline partial,” while the source provides no hardware, prompts, software versions, test dates, sampling settings, scoring process or evidence that the results were reproduced. Those differences and missing details mean the figures cannot be treated as a verified or necessarily comparable cross-model advantage.

The comparison image is evidence of what was displayed in the leak, not independent evidence that the figures are accurate.

Displayed context-window and token-price figures

A context window is the token limit for the context a model can process in a request or conversation. Depending on the system, that limit may include some or all of the generated output as well as the conversation or other input context. The image labels Gemini 4 Pro’s context window as 2 million tokens, compared with 1 million tokens for Claude Opus 5.5 and 1.1 million tokens for GPT-6 Astra. The supplied thread later says that Gemini 4 Pro is “confirmed to be 1M context not 2M,” so the alleged Gemini figure remains unresolved.

The image also displays these prices per 1 million tokens:

  • Gemini 4 Pro: $2.25 / $11.25

  • Claude Opus 5.5: $4 / $20

  • GPT-6 Astra: $10 / $50

The image does not explain what the two prices in each pair represent. It does not identify input and output billing, service tiers, caching, batch discounts, regional terms or whether the figures are proposed, unofficial or final. Readers should not treat the displayed amounts as confirmed API pricing.

Neither the image’s 2-million-token label nor the thread’s 1-million-token correction is independently verified by the supplied source.

What the comparison does—and does not—show

If the displayed figures are accurate and comparable, they would suggest a potentially strong benchmark and price position for the alleged Gemini model. They do not show how the models perform on a particular user’s workload, how quickly they generate responses, or what their effective cost would be in production.

A benchmark score can depend on the task definition, evaluation harness, prompts and scoring rules. A displayed token price also cannot be compared reliably without knowing which type of tokens each number covers and whether the services use equivalent conditions. The thread’s predictions about competitive pricing, margins, speed and market effects remain speculation.

Open questions about the leak

The leak leaves several material questions unanswered:

  • Has Google officially announced Gemini 4 Pro or published documentation for it?

  • Who produced the comparison and under what test conditions?

  • Were the benchmark results independently reproduced?

  • Are the model names in the image consistent with the names used in the thread?

  • Is the context window 2 million tokens, 1 million tokens or something else?

  • Do the displayed prices represent official, final API rates, and what do the slash-separated figures mean?

  • Is the alleged model available to users, and if so, through which service?

Until those questions have answers from reliable documentation or reproducible testing, the comparison is useful mainly as an account of what the leak is reporting. It should not be treated as confirmation that Gemini 4 Pro has launched, definitively beats Claude or OpenAI models, or will force competitors to change their prices.

Sources