Kimi K2.8 Preview is fully rolled out in Kimi Code, according to kimi.com’s “What’s New” page entry dated September 11, 2026. Kimi says the release keeps the kimi-for-coding model ID, adds adjustable thinking effort and has a context window of up to 1 million tokens across all membership tiers.
Source: What's New · kimi.com
The announcement positions K2.8 Preview as a model with performance close to K3 and more efficient thinking than K2.7 Code. Those performance and efficiency statements come from Kimi’s release notes; the supplied source does not include benchmark results, testing methodology or independent evaluation.
Existing Kimi Code clients do not need a model-ID migration
The most immediately practical part of the release is that K2.8 Preview uses the existing kimi-for-coding model ID. Clients and third-party tools already configured to use that identifier should continue making requests without a model-setting change, according to Kimi.
That makes the rollout different from a migration that requires users to replace an old model name throughout API clients, scripts or integrations. The announcement supports the unchanged identifier and no-configuration-change point, but it does not document every provider-specific compatibility detail.
Adjustable thinking effort gives users more control
K2.8 Preview supports the same three thinking-effort levels listed for K3: low, high and max. The release notes say that max is the default.
Thinking effort controls how much reasoning work the model applies before producing an answer. In practical terms, the available settings let users choose between a lower-effort response and a more intensive mode, although the supplied announcement does not describe token costs, latency differences or recommended settings for particular tasks.
Kimi also says that when thinking is turned off, requests to the K3 series and K2.8 Preview are served by K2.8 Preview in no-thinking mode. This gives users a separate option from the low, high and max settings, but the release notes do not provide detailed configuration instructions for each client or provider.
The 1-million-token context window is available across membership tiers
Kimi says K2.8 Preview supports a context window of up to 1 million tokens across all membership tiers. A context window is the amount of conversation, code and other material a model can consider in a request. A larger window can be useful when working with an extensive codebase or a long-running task because more relevant material can remain available at once.
The announcement does not specify additional quotas, rate limits, pricing or how much of the maximum context a particular request will consume. The 1-million-token figure should therefore be read as the stated maximum context availability, not as a guarantee that every request or workflow will use that much context.
How K2.8 Preview is positioned against K2.7 Code and K3
The release notes describe K2.8 Preview as having more efficient thinking than K2.7 Code and performance close to K3. They do not provide a benchmark table or independent test results to show how large either difference is.
Model or release | Thinking effort described in the source | Context information described in the source | Positioning in the release notes |
|---|---|---|---|
Kimi K2.8 Preview | Low, high and max; max is the default. No-thinking requests are served by K2.8 Preview. | Up to 1 million tokens across all membership tiers | Kimi describes performance as close to K3 and thinking as more efficient than K2.7 Code. |
K2.7 Code | The supplied release excerpt does not specify its effort levels. | The supplied excerpt does not specify its context availability. | Used as the comparison point for K2.8 Preview’s reported thinking efficiency. |
K3 | Low, high and max | Up to 1 million tokens for Allegretto members and above, according to the supplied release notes | Used as the performance reference for K2.8 Preview; no independent comparison data is supplied. |
The comparison therefore describes Kimi’s intended positioning rather than establishing that K2.8 Preview is universally better than K3. Users evaluating the models should treat the differences in task quality, speed and resource use as questions for their own workloads unless they have separate test evidence.
What the announcement establishes for users
Kimi’s release entry establishes the reported rollout, model-ID continuity, thinking controls and stated context availability. Before relying on those capabilities, users should check which controls their particular Kimi Code client exposes and how that client handles context limits. The supplied page extract does not provide the client-specific configuration, quota or pricing details needed to answer those questions.





0 comments
No approved comments yet. You can start the conversation.
Leave a comment