In a post on X, @chetaslua reports that GPT-6 shows lower community-measured “juice” values in Codex than in the API across the listed effort levels. The attached graphic reports the largest maximum-effort gaps for GPT-6 Astra, Sol and Luna, while also stating that the measurements are not confirmed by OpenAI.

Image shared by @chetaslua on X
Image shared by @chetaslua on X

Image credit: @chetaslua on X

Here, “juice” refers to the reasoning budget that the model reportedly reveals through a community “juice prompt,” according to the graphic. The evidence does not establish that Codex uses a different model, produces lower-quality answers or has a universally lower limit. It shows only the values reported in this one source.

The reported API and Codex values

The graphic lists API and Codex values for five effort levels: low, medium, high, xhigh and max. The figures below reproduce those reported values. A question mark indicates that the graphic does not provide a medium-effort Codex measurement for GPT-6 Luna.

GPT-6 variant

Low: API / Codex

Medium: API / Codex

High: API / Codex

Xhigh: API / Codex

Max: API / Codex

Astra

4 / 2

10 / 4

24 / 10

64 / 32

832 / 128

Sol

4 / 1

12 / 10

24 / 15

64 / 21

768 / 62

Luna

4 / 1

12 / ?

24 / 6

64 / 25

768 / 70

Values reported in the attached community graphic. They are described there as community-measured and not confirmed by OpenAI; the Luna medium-effort Codex value was not measured.

The API value is higher than the Codex value at every listed effort level where both values are available. The gap is relatively small for some lower-effort Sol measurements: the graphic lists 12 for the API and 10 for Codex at medium effort. Other entries show a larger separation, such as Astra’s 64 versus 32 at xhigh effort and Luna’s 64 versus 25.

What the maximum-effort comparisons show

At maximum effort, the graphic reports these API-versus-Codex pairs:

  • GPT-6 Astra: 832 in the API versus 128 in Codex, or about 6.5 times as much reported juice in the API.

  • GPT-6 Sol: 768 in the API versus 62 in Codex, or about 12.4 times as much reported juice in the API.

  • GPT-6 Luna: 768 in the API versus 70 in Codex, or about 11 times as much reported juice in the API.

These ratios describe the displayed or inferred budget values in the graphic. They do not measure answer quality, coding ability, latency, cost, reliability or the amount of useful reasoning produced in practice. A larger reported budget also does not by itself prove that one configuration performs better on a particular task.

What the report does not establish

The supplied evidence is one X post and its attached graphic. It does not describe the full testing procedure, the number of runs, the exact prompt beyond the reference to a community “juice prompt” or how the displayed values were inferred.

It also does not establish whether the API and Codex tests used identical model versions, settings or effort controls. The post’s “same GPT-6” framing is therefore not an independently verified finding. The graphic’s wording supports reporting the values as community measurements, but not presenting them as OpenAI documentation or guaranteed product limits.

Finally, the report does not connect the budget differences with a difference in output quality or coding performance. Readers can reasonably conclude that this source reports lower observed or inferred Codex juice values for the three listed GPT-6 variants. They cannot conclude from it alone why the values differ, whether the pattern applies beyond these measurements or whether OpenAI officially supports the reported numbers.

Sources