xAI describes Grok 4.7 as its most capable model for coding and knowledge work, with a focus on tasks that run for hours rather than short, isolated prompts. The announcement says the model uses a larger base model than Grok 4.6, underwent a longer reinforcement-learning training run, and is better at checking its own work and managing longer context. xAI also lists several access routes.
The results and safety figures below are reported by xAI. The announcement does not provide independent testing, detailed methodologies, sample sizes, confidence intervals, or evidence that the results will transfer to every user’s workload.
What Grok 4.7 is designed to do
Grok 4.7 is intended for longer-running coding, terminal, document, presentation, legal, clinical, and engineering tasks. xAI says it trained the model on a harder mix of problems weighted toward tasks that take many hours to complete. It also says Grok 4.7 was trained to natively understand the Grok Bot harness, which is meant to improve conversational tasks and broader knowledge work.
The practical emphasis is therefore less about a single new feature and more about persistence: continuing through a complex task, reviewing intermediate work, and handling a larger amount of context. Those are design goals described by xAI, not a guarantee that the model will complete every long-running task successfully.

How Grok 4.7 compares with Grok 4.6 and other models
xAI’s announcement compares Grok 4.7 with Grok 4.6 and two other named models across coding, office work, terminal tasks, legal work, clinical reasoning, and electrical engineering. The published figures are reported results rather than a universal ranking. The source does not explain every benchmark’s scoring scale or provide the full test conditions.
Benchmark or task area | Grok 4.7 (xhigh) | Grok 4.6 (high) | GPT-5.6 Sol (max) | Fable 5.1 (max) |
|---|---|---|---|---|
CursorBench 4.0, software engineering | 46.3% | 40.4% | 41.7% | 51.8% |
DeepSWE v1.1, software engineering | 71.0%* | 65.2% | 72.7% | 70.0% |
AA Briefcase v1.1, multi-hour office work | 1,657 | 1,546 | 1,487 | 1,678 |
Terminal-Bench 4.0, multi-hour terminal work | 38.0% | 20.3% | 37.3% | 57.9% |
Harvey Legal Agent Benchmark | 19.6% | 15.8% | 2.5% | 6.7% |
HealthBench Professional | 56.7% | 48.5% | 60.5% | 62.1% |
EEBench, electrical engineering | 64.0% | 53.0% | 39.4% | 56.4% |
*The asterisk marks Grok 4.7’s high-effort DeepSWE result, according to xAI.
The xhigh and high labels indicate the listed effort settings for Grok 4.7 and Grok 4.6. Those settings matter when interpreting both performance and price comparisons: a result may involve a different amount of model effort, and the table should not automatically be read as a comparison of identical configurations.
On the reported figures, Grok 4.7 improves on Grok 4.6 in every listed comparison. Among the percentage-based comparisons, its largest increase is on Terminal-Bench 4.0, rising from 20.3% to 38.0%. That percentage-point change should not be compared directly with the raw-score change on AA Briefcase, which uses a different scale.
The figures also do not make Grok 4.7 the top model in every category shown. Fable 5.1 has the highest listed scores on CursorBench, AA Briefcase, Terminal-Bench, and HealthBench Professional. GPT-5.6 Sol is higher on DeepSWE and HealthBench Professional, while Grok 4.7 leads the listed Harvey Legal Agent Benchmark and EEBench comparisons.
Coding and terminal work
The coding results support xAI’s emphasis on longer tasks, but the benchmarks evaluate different kinds of work. CursorBench 4.0 is presented as a test of longer-running software-engineering tasks, while Terminal-Bench 4.0 focuses on multi-hour terminal work. A strong result on one does not necessarily predict the same result on the other.
DeepSWE v1.1 adds another qualification because Grok 4.7’s listed score is marked as high effort. Without the benchmark methodology and effort settings for every model, these results are best read as xAI’s reported comparison rather than as a complete independent performance test.

Documents and professional knowledge work
xAI says Grok 4.7 is better at creating documents and presentations. It cites GDPval and AA Briefcase as evaluations involving tasks performed by professionals such as lawyers, nurses, and financial analysts.
The announcement says Grok 4.7 improves on Grok 4.6 on both benchmarks and performs comparably with other frontier models. For a potential user, that supports an intended use case around multi-step professional work, not a verified claim that Grok 4.7 can replace a lawyer, nurse, financial analyst, or other professional. Human review remains important for consequential documents and decisions.
Safety, refusals, and cybersecurity
xAI says Grok 4.7 uses a new safeguard stack and is the strongest model it has tested for refusal behavior and jailbreak resistance. In biological and cybersecurity scenarios, the company presents the model as aiming to preserve usefulness for benign tasks while refusing dangerous requests.
The announcement reports a 62.4% result on LatchBio’s biosafety benchmark and says Grok 4.7 topped that evaluation. It also reports that the model allowed 3.3% of risky dual-use prompts through on HackerBench v0.3, which xAI describes as its benchmark for risky and malicious cyber tasks. The company says the model rarely blocks legitimate security work and has begun offering selected cybersecurity partners invite-only access to its red-team capabilities for defense research.
These are company-reported safety results, not a complete assessment of the model’s safeguards. The announcement does not describe the benchmark samples, scoring procedures, or the definition of legitimate security work in enough detail to independently interpret the figures. The 3.3% figure also should not be read as a general error rate for all cybersecurity prompts.
Grok 4.7 pricing and availability
xAI lists Grok 4.7 at a starting price of $2 per million input tokens and $6 per million output tokens. It also says a fast variant provides twice the output speed at twice the price. The announcement does not state the fast variant’s exact output speed or provide API rate limits, billing terms, context-window details, regional availability, or service-level commitments.
The source lists Grok 4.7 as available through:
Cursor
Grok Build
The Grok API
Third-party coding harnesses
Model routers
Cloud platforms
xAI also says users can try Grok Build for free. Access-channel details may differ, including the available variant, limits, pricing, and performance, so users should check the terms for the route they plan to use.
The announcement presents an apparent pricing inconsistency. Its opening description says Grok 4.7 is served at the same price and speed as Grok 4.6, while the comparison table lists Grok 4.7 at $2 per million input tokens and $6 per million output tokens, versus $4 and $20 for Grok 4.6. The table therefore shows lower listed token prices for Grok 4.7, but the different xhigh and high labels mean the figures should not automatically be treated as identical configurations.
For API buyers, input and output token prices are only part of the decision. Total cost will also depend on how much context a task uses, how many attempts or verification passes it requires, whether the fast variant is selected, and the limits imposed by the chosen access channel. The announcement does not provide enough information to calculate a typical task cost.

What the announcement does—and does not—show
Grok 4.7 is presented as a model for extended coding and knowledge-work tasks, with a reported improvement over Grok 4.6 in the comparisons shown. xAI also positions it as a model that can combine stronger performance on legitimate cybersecurity work with more selective refusals for dangerous requests.
The announcement does not include independent hands-on testing, benchmark methodology, real-world examples, or results from typical user workloads. Readers evaluating Grok 4.7 should therefore treat the reported scores as signals about xAI’s intended positioning, then check the relevant access channel’s pricing, limits, and performance on their own tasks before committing to it.





0 comments
No approved comments yet. You can start the conversation.
Leave a comment