September 9, 2026:


The most-watched independent AI benchmark in the industry revised its scoring system three times in under a week, producing a different #1 ranking each time — and the score a developer saw on Monday was not the score in place by Sunday. The latest version, Intelligence Index v4.3, published September 7, lands OpenAI’s GPT-6 Astra and Anthropic’s Claude Fable 5.1 at an identical 53 points apiece. The tie is real. The choice between the two models is not.
Artificial Analysis, the independent benchmarking firm that operates the Intelligence Index, was explicit that the three-version sprint — v4.1.1 on September 3, v4.2 on September 4, v4.3 on September 7 — was an accelerated rollout of changes planned for a larger Intelligence Index v5 update. The v4.2 release rationale stated that the frontier had moved fast enough in recent weeks to make waiting untenable. That explanation is credible. It is also precisely the situation in which a benchmark’s live-document nature produces consequences its authors did not intend: any production decision made using last week’s leaderboard was made on a methodology that no longer applies.
The full arc of the revision sequence shows how much a composite benchmark’s composition governs its conclusions. Under Intelligence Index v4.1.1 — the version in place when GPT-6 Astra launched on September 3 — Astra scored 61, tied with its predecessor GPT-5.6 Sol and five points behind Fable 5.1 at 66. A model OpenAI described as a generational leap scored identically to the model it replaced on the leading third-party composite benchmark. The Astra v4.1.1 benchmarking results document this gap in detail.
Four days later, v4.2 incorporated AA-Briefcase (an in-house agentic knowledge-work evaluation with a private test set) and GDP.pdf (long-context document reasoning across 4,592 PDF pages), while removing GPQA Diamond on the grounds that frontier models had saturated it. The v4.2 evaluation changes described the full changelog. The gap narrowed: Fable 5.1 at 57, Astra at 55. Three days after that, v4.3 upgraded Terminal-Bench from v2.1 to v4.0 and replaced τ³-Banking — a customer support automation benchmark — with AutomationBench-AA, a 657-task business workflow benchmark covering Finance, HR, Marketing, Operations, Sales, and Support. The gap closed entirely to 53-53.
The trajectory — 66 vs. 61, then 57 vs. 55, then 53 vs. 53 — is not evidence that the models converged. No new models were released. The benchmarks changed.
Identical composite scores mask two findings that matter more for deployment than the headline number does.
The first is a cost gap. The Astra v4.3 cost comparison shows that at the same 53-point intelligence score, GPT-6 Astra costs $3.26 per Intelligence Index task on a weighted-average basis, while Claude Fable 5.1 costs $7.63 — a 57% lower cost per task for Astra at equivalent composite performance. The mechanism behind the gap is token efficiency: Astra consumed approximately 60 million output tokens to complete the full Intelligence Index evaluation suite, while Fable 5.1 consumed approximately 190 million — more than three times as many tokens to produce the same composite score. At identical base API prices of $10 per million input tokens and $50 per million output tokens, higher output token consumption translates directly into higher cost per task.
The second is sub-benchmark divergence. The tie is an aggregation outcome, not a statement of general equivalence. According to the v4.3 sub-benchmark score breakdown, Fable 5.1 outperforms Astra on AA-Briefcase (the private agentic knowledge-work evaluation) and SciCode (Python-based scientific computing problems). Astra leads on Terminal-Bench v4.0, where it scored 59.1% against Fable 5.1’s 52.0%, and on AutomationBench-AA, where Astra leads the field at 68.5%.
The practical split: Astra is the stronger option for terminal-based software work and multi-application business workflow automation. Fable 5.1 is the stronger option for long-horizon knowledge work and scientific computing.
Both new evaluations introduced in v4.3 are materially different from what they replaced — and understanding the mechanics explains why the scores shifted as they did.
The Terminal-Bench v4.0 evaluation details show that the benchmark tests whether an AI agent can complete complex, multi-step work using only a terminal interface. The evaluation spans seven domains: software engineering, machine learning, science, operations, security, hardware, and media. Artificial Analysis runs all 66 tasks three times each and reports average pass@1 — the percentage of tasks the model completes correctly on the first attempt. The upgrade from Terminal-Bench v2.1 recalibrated both compute and time allowances and tightened task verification. On this harder benchmark, Astra’s 59.1% is 19.2 percentage points ahead of its predecessor Sol’s 39.9% — a meaningful capability jump within the OpenAI family that the prior index versions did not capture. Claude Opus 5 scored 49.0% on the same test.
The AutomationBench-AA evaluation overview, developed in collaboration with Zapier, covers 657 tasks across six business function categories. Agents must discover the relevant APIs for each task independently rather than being pre-configured — a design choice that tests real-world adaptability rather than scripted tool-calling. Artificial Analysis awards partial credit for completed objectives, but any guardrail violation on a task resets that task’s score to zero. The benchmark uses a private held-out test set, meaning labs cannot train or optimize to the specific questions.
On the clean-completion rate comparison — the share of workflows where every objective is completed without any guardrail violation — Astra completed 41.6% of workflows cleanly, against 32.1% for Fable 5.1 and 28.3% for Claude Opus 5. Completing every objective while respecting all guardrails is substantially harder than completing part of a workflow, which is why the headline score and the clean-completion score diverge by roughly nine percentage points for each model.
The revision to private held-out test sets is the structural change that matters most for long-run benchmark reliability. In v4.2, 40% of the Intelligence Index’s weighting came from private test sets — up from 20% in v4.1. In v4.3, that figure rose to 45%. Artificial Analysis has indicated it will continue to increase this percentage in Intelligence Index v5. The private test set methodology documents the full weighting table.
The reason is Goodhart’s Law applied to AI benchmarking: once a test becomes a target, it ceases to be a good test. When benchmark questions are public, AI labs can — and do — train models on datasets that overlap with those questions, producing scores that reflect optimization for the benchmark rather than genuine capability. GPQA Diamond, removed in v4.2, was an example: Artificial Analysis cited GPQA Diamond saturation removal — frontier models scoring near-perfect — as the reason for its removal. Saturation is often Goodhart’s Law arriving at its endpoint.
Private test sets raise the cost of gaming without eliminating it — labs can still probe a private benchmark iteratively through API calls and infer its structure — but they meaningfully reduce the reliability of any specific optimization strategy. The tradeoff is transparency: a benchmark that cannot be independently reproduced by the community must be taken on trust. As the private-weighting percentage rises, Artificial Analysis’s methodology documentation and audit credibility become more important, not less.
Below the Astra/Fable 5.1 tie, v4.3 places Claude Opus 5 (max) at 51, Claude Fable 5 (with fallback) at 50, Muse Spark 1.3 (max) at 48, and GPT-5.6 Sol (max) at 47. The full v4.3 leaderboard scores are available on Artificial Analysis’s evaluation page.
In the open-weights tier, GLM-5.3 Flash scores 42, Qwen3.8 2.4T A95B scores 40, and DeepSeek V4 Pro 0813 (max) scores 36.
The cost-efficiency frontier is currently occupied by four labs. OpenAI dominates it, with all five reasoning-effort configurations of GPT-6 Astra offering the lowest cost at their respective intelligence levels. MiMo-V2.5-Pro, GLM-5.3 Flash, and Claude Fable 5.1 at extra-high effort (max) round out the Pareto frontier. At their respective intelligence-level thresholds, GLM-5.3 Flash at $0.25 per task costs 18% of what GPT-5.6 Terra costs at the same 42-point intelligence score.
Artificial Analysis’s three-version sprint is not a benchmark failure — it is benchmark evolution running at an accelerated pace because the underlying models are developing faster than the evaluation methodology can comfortably track. That is, in fact, the situation the benchmarking community faces across the industry, and Artificial Analysis’s transparency about the revision rationale is more editorial honesty than many benchmark providers offer.
What the revision pace does mean, practically, is that any production model selection decision that references an Artificial Analysis Intelligence Index score should specify the version. A score without a version number is a claim of uncertain vintage; the number that mattered on September 3 was different from the number that matters on September 9. The organization has said it plans further incremental updates before Intelligence Index v5. The current 53-53 tie is accurate as of v4.3, published September 7, 2026 — and it is exactly as provisional as every prior version was.
For developers choosing between Astra and Fable 5.1 today: equal composite scores, 57% lower per-task cost, and Astra’s structural advantage in terminal-based engineering and multi-application workflow automation together make Astra the cleaner choice for teams whose primary workloads sit in those categories. For long-horizon knowledge work and scientific computing, Fable 5.1’s sub-benchmark advantages persist beneath the headline tie.
Not exactly. The revisions reflect a genuine methodological challenge: as frontier models approach or exceed the difficulty ceiling of existing benchmarks, evaluators must raise the bar or the benchmark becomes uninformative. Artificial Analysis removed GPQA Diamond in v4.2 precisely because frontier models had saturated it — scoring near-perfect. Adding harder, more realistic tasks (Terminal-Bench v4.0, AutomationBench-AA) and more private test sets is the structurally correct response, even if the pace of revision is disorienting. The more useful adjustment for users is to treat any Intelligence Index score as version-specific rather than as an absolute number, and to check whether the version underlying a cited score matches the current methodology before using it for a production decision.
Because a composite score is an average, not a specification. Two models can produce the same average score on different sub-benchmarks — as Astra and Fable 5.1 do — and then generate very different token outputs in the process. Fable 5.1 used approximately 190 million output tokens to complete Artificial Analysis’s full evaluation suite; Astra used approximately 60 million. At $50 per million output tokens, that difference alone accounts for a large portion of the $7.63 versus $3.26 per-task gap. For any deployment at scale, a 57% cost differential at identical aggregate performance is a material selection factor.
τ³-Banking, the benchmark AutomationBench-AA replaced, tested customer support automation in a single domain — financial services customer contacts. AutomationBench-AA covers 657 tasks across six business function categories, uses a private held-out test set to reduce gaming, requires agents to discover APIs independently rather than using pre-configured tools, and penalizes guardrail violations by zeroing any task where they occur. The broader coverage and the guardrail-violation penalty make it a substantially more realistic test of enterprise AI deployment readiness than its predecessor.
Yes, explicitly. Artificial Analysis stated in the v4.3 release that the v4.2 and v4.3 changes are a continuation of its v5 rollout and that further incremental updates are planned before v5 arrives. The private test set percentage is expected to increase further. Developers relying on the Intelligence Index for production model selection should monitor version announcements rather than treating any current number as stable.