Benchmarking firm Artificial Analysis released version 4.2 of its Intelligence Index on September 4, 2026, an interim update that reworks how frontier AI models are measured — and, according to the new leaderboard, Anthropic's Claude Fable 5.1 holds the top spot, with OpenAI's GPT-6 Astra in second place after a four-point gain over GPT-5.6 Sol.

The update comes just eight months after Index v4 launched in January, an unusually quick turnaround the firm attributes to the pace of recent frontier releases. "With the frontier moving so quickly in the past weeks, we feel it is important to deliver an immediate interim update to ensure our Index remains as relevant and useful as ever," the company wrote in its announcement.

The headline rankings

According to the published results, Anthropic's Claude Fable 5.1 leads the Index, followed by OpenAI's GPT-6 Astra, which improved four points over GPT-5.6 Sol. Meta ranks as the third-strongest lab on the leaderboard, followed by SpaceXAI, Moonshot/Kimi, Z.AI, and Google.

Anthropic and OpenAI also share the updated cost-efficiency picture: the two companies, along with Meta and Z.AI, occupy the Index's "Cost per Task" Pareto frontier, meaning no other lab currently offers strictly better intelligence for the same price per completed task.

On token efficiency, Artificial Analysis found GPT-6 Astra "more token efficient than almost every other model near the intelligence frontier," with Claude Fable 5.1, Grok 4.5, and Gemini 3.5 Flash-Lite sitting at either end of the curve. The firm noted in a companion analysis, however, that Astra's lower token usage is outweighed by its higher prices — a reminder that raw efficiency and total cost do not always move together.

What changed in v4.2

The most significant structural change is a deliberate push against benchmark gaming. Forty percent of the Index's weighting now comes from private, held-out test sets — double the share in v4.1 — including AA-Briefcase, AA-Omniscience, and solutions for CritPt. The firm said that figure will rise further in Index v5.

Two new evaluations anchor the update:

  • AA-Briefcase, an in-house agentic knowledge work benchmark with a private held-out test set. Models are evaluated on multi-week professional projects, each with many linked tasks and thousands of input source files, and graded through a combination of rubric scoring and pairwise comparison covering verifiable task success, analytical quality, and presentation quality.
  • GDP.pdf, created by Surge AI, which tests single-turn reasoning across professional documents: 100 PDFs spanning ten domains, with evidence distributed across 4,592 pages of text, tables, charts, and footnotes. Responses are graded against 1,275 expert-authored atomic criteria, and the headline All-pass Rate credits a task only when every criterion is satisfied.

One long-standing benchmark is on the way out: GPQA Diamond, the graduate-level science reasoning test, has been removed because it has been effectively saturated by frontier models.

The firm also upgraded its grading infrastructure, releasing AA-LCR v1.1 with a new grading system prompt and corrected answer keys, re-anchoring the Elo scale for GDPval-AA v2 and AA-Briefcase so ratings stay stable as new models arrive, and hardening SciCode's grading sandboxes so that slow-but-correct code no longer counts as a failure.

What the new tests reveal

The results on the new evaluations offer a more granular picture of where the frontier labs actually stand.

On AA-Briefcase, Anthropic's Claude Fable 5.1 and Opus 5 lead, followed by OpenAI's GPT-6 Astra and Meta's Muse Spark 1.3. GPT-6 Astra scores roughly 85 Elo points above GPT-5.6 Sol on the same evaluation — a substantial jump between successive OpenAI generations.

On GDP.pdf, the picture flips: OpenAI leads with GPT-6 Astra at 33.2% All-pass Rate, followed by GPT-5.6 Sol at 28.2% and Claude Fable 5.1 at 26.2%. The divergence suggests long-document synthesis over very large contexts remains a relative strength for OpenAI's newest model, even as Anthropic's models lead the overall Index and agentic work tasks.

Why private test sets matter

The shift toward held-out evaluation reflects a broader credibility problem in AI benchmarking. As models have absorbed more public test data, labs and independent observers have grown increasingly concerned that headline scores on well-known benchmarks reflect memorization or targeted training rather than genuine capability. By keeping a growing share of its test sets private and weighting them more heavily, Artificial Analysis is attempting to preserve the signal that public leaderboards have been losing.

The firm said it has been building elements of Index v5 "for months" and deliberately held back updates to keep the Index stable through the recent wave of major model launches. Beyond the interim v4.2 release, it plans more incremental updates in the near future as it completes the v5 transition.

For buyers choosing between frontier models, the practical takeaway from v4.2 is that the top of the market remains a two-horse race between Anthropic and OpenAI, with Meta closing the gap — and that the evaluations designed to be resistant to gaming are now doing more of the ranking work. Users tracking model selection decisions can follow the latest AI developments as the v5 rollout approaches.

Stay Ahead of AI

Benchmark updates like this one reshape how the industry judges progress. Read more AI news and stay ahead of every model release.

Read more AI news →