News Report Technology
September 22, 2026

xAI Ships Grok 4.7 With New Safeguard Stack: Independent Benchmarks Confirm Gains, Flag Doubled Token Consumption

In Brief

xAI launches Grok 4.7 with frontier coding and agentic work performance, a new safety stack, and $2/$6 per million token pricing. Independent benchmarks confirm the gains but flag doubled token use.

xAI Ships Grok 4.7 With New Safeguard Stack: Independent Benchmarks Confirm Gains, Flag Doubled Token Consumption

xAI has released Grok 4.7, its most capable model to date for coding and knowledge work, positioning it as a frontier offering at a competitive price point. The model is built on a new, larger base than its predecessor, Grok 4.6, and was trained with an extended reinforcement learning run on a harder mix of tasks weighted toward problems that require many hours to complete. According to the company, these improvements enhance the model’s code generation, long-context handling, and self-verification capabilities, while native understanding of the Grok Bot harness makes it more effective at conversational tasks and general knowledge work.

Benchmark results illustrate where the gains are most pronounced. On CursorBench 4.0, which stresses longer-running coding tasks, Grok 4.7 posted a score of 46.3%, up from 40.4% on Grok 4.6 and ahead of GPT-5.6 Sol’s 41.7%, while remaining behind Fable 5.1’s 51.8%. The model also recorded a high-effort score of 71.0% on DeepSWE v1.1, trailing GPT-5.6 Sol’s 72.7% but edging Fable 5.1’s 70.0%. 

In professional knowledge work, Grok 4.7 improved on its predecessor across GDPval and AA Briefcase, the latter of which evaluates multi-hour office tasks performed by professionals such as lawyers, nurses, and financial analysts, where it scored 1,657, close behind Fable 5.1’s 1,678. The most striking jump came on Terminal-Bench 4.0, a measure of multi-hour terminal work, where the model nearly doubled Grok 4.6’s performance, rising from 20.3 to 38.0%. On HackerBench v0.3, which covers risky cyber tasks, the model allowed only 3.3% of dangerous dual-use prompts through while rarely blocking legitimate security work.

A distinguishing element of the release is Grok 4.7’s safety stack, described by the company as entirely new and the strongest it has tested on refusals and jailbreak resistance. The model leads on LatchBio’s biosafety benchmark at 62.4%, balancing utility for benign tasks with safe refusal on dangerous ones in dual-use domains such as cybersecurity and biological research. xAI has also begun granting select cybersecurity partners invite-only access to the model’s red-team capabilities for defense research.

Pricing is a central part of the launch strategy. Grok 4.7 is served at the same rates as Grok 4.6, two dollars per million input tokens and six dollars per million output tokens, undercutting GPT-5.6 Sol (four and twenty dollars respectively) and Fable 5.1 (ten and fifty dollars). A fast variant offering twice the output speed is available at twice the price. On CursorBench’s price-performance frontier, this places Grok 4.7 among the most cost-efficient options for extended coding workloads.

The model is available immediately in Cursor and Grok Build, with access also provided through the Grok API, third-party coding harnesses, and model routers and cloud platforms. The release intensifies competition in the frontier model segment, where vendors are increasingly differentiating on long-duration agentic tasks, safety calibration, and cost per completed task rather than raw benchmark scores alone. For developers and enterprises evaluating coding-oriented models, Grok 4.7 presents a combination of pricing stability, improved multi-hour task performance, and hardened safeguards that could make it a viable alternative to more expensive frontier offerings.

Independent Analysis Confirms Frontier Status, With a Token-Consumption Caveat

Independent benchmarking firm Artificial Analysis has corroborated xAI’s claims, scoring Grok 4.7 at 46 on its Intelligence Index, a two-point gain over Grok 4.6, a result that places SpaceXAI among the top four AI labs. The firm’s evaluation, conducted at xhigh reasoning effort, found the model’s clearest advance in long-horizon agentic knowledge work: on its private AA-Briefcase benchmark, Grok 4.7 gained 111 Elo over its predecessor to reach 1,657, ranking alongside Claude Opus 5 and Claude Fable 5.1 at the frontier. The improvement was driven primarily by analytical quality, which rose sharply to 1,994 Elo from 1,690, while presentation quality held roughly steady. On GDPval-AA, which measures practical work products such as documents, spreadsheets, and slides, the model scored 1,695 Elo, up 90 points.

Performance gains came at a cost in compute consumption. Artificial Analysis measured approximately 81,000 output tokens per Intelligence Index task for Grok 4.7 at xhigh, more than double the 36,000 used by Grok 4.6 and nearly triple the 27,000 consumed by GPT-6 Astra at max effort. The firm also noted incremental changes elsewhere: modest improvements on Terminal-Bench 4.0 and GDP.pdf, alongside small regressions on AA-LCR and AutomationBench-AA. Reliability metrics improved modestly, with the AA-Omniscience hallucination rate falling to 29% from 34%, while accuracy held near-unchanged at 47%. Technical specifications remain consistent with the previous generation, including an unchanged 500,000-token context window and cache-hit pricing discounted to $0.50 per million tokens.

Disclaimer

In line with the Trust Project guidelines, please note that the information provided on this page is not intended to be and should not be interpreted as legal, tax, investment, financial, or any other form of advice. It is important to only invest what you can afford to lose and to seek independent financial advice if you have any doubts. For further information, we suggest referring to the terms and conditions as well as the help and support pages provided by the issuer or advertiser. MetaversePost is committed to accurate, unbiased reporting, but market conditions are subject to change without notice.

About The Author

Alisa, a dedicated journalist at the MPost, specializes in crypto, AI, investments, and the expansive realm of Web3. With a keen eye for emerging trends and technologies, she delivers comprehensive coverage to inform and engage readers in the ever-evolving landscape of digital finance.

More articles
Alisa Davidson
Alisa Davidson

Alisa, a dedicated journalist at the MPost, specializes in crypto, AI, investments, and the expansive realm of Web3. With a keen eye for emerging trends and technologies, she delivers comprehensive coverage to inform and engage readers in the ever-evolving landscape of digital finance.

Hot Stories
Join Our Newsletter.
Latest News

Shufti, Jumio, Sumsub, And Beyond: Top 6 Identity Verification And Compliance Platforms To Know In 2026

Shufti, Sumsub, Incode, Veriff, Persona and Jumio compared on compliance lifecycle coverage, pricing transparency and fraud detection ...

Know More

2026 AI Market Claims Vs SEC Fillings: Linkmate Analysis

Is the AI market really all just PR talk or there's a deeper math going on in ...

Know More
Read More
Read more
Arthur Hayes: AI ‘Safety’ Slowdown May Expose Trillion-Dollar Debt Bubble, With Government Backstop Set To Boost Bitcoin
Markets News Report Technology
Arthur Hayes: AI ‘Safety’ Slowdown May Expose Trillion-Dollar Debt Bubble, With Government Backstop Set To Boost Bitcoin
September 22, 2026
Agora, Catena, And Bastion Win Conditional OCC Nods As Federal Trust Charter Wave Accelerates
Business News Report Technology
Agora, Catena, And Bastion Win Conditional OCC Nods As Federal Trust Charter Wave Accelerates
September 22, 2026
Gate Update: Record Inflows, Bali Summit, And A New Platform Vision On The Horizon
Digest News Report Technology
Gate Update: Record Inflows, Bali Summit, And A New Platform Vision On The Horizon
September 21, 2026
AIBC World 2026: Rome Gears Up For Europe’s Frontier Technology Summit This November
Lifestyle News Report Technology
AIBC World 2026: Rome Gears Up For Europe’s Frontier Technology Summit This November
September 21, 2026