Featured image of post Artificial Analysis Intelligence Index v4.2 Released: Private Test Sets and Real-World Tasks Promote Benchmark Maturity

Artificial Analysis Intelligence Index v4.2 Released: Private Test Sets and Real-World Tasks Promote Benchmark Maturity

v4.2 raises private test weight to 40%

Core Announcement

Core Announcement
Core Announcement|News screenshot

Artificial Analysis has released Intelligence Index v4.2, an interim update ahead of its upcoming v5 release. The company says the update is intended to keep pace with fast-moving frontier models, make the Index more relevant to real-world use cases, and reduce benchmark gaming through more private held-out test sets. Key facts:

  • Version: v4.2, an interim update to v4 while v5 remains in development
  • Main changes: Adds AA-Briefcase and GDP.pdf; removes GPQA Diamond after saturation
  • Private test weighting: 40% of the Index weighting now comes from private held-out test sets, double the share in v4.1
  • Next steps: Artificial Analysis says the held-out percentage will rise further in Index v5, with more incremental releases planned in the near future

New Tasks: More Realistic Knowledge Work

New Tasks: More Realistic Knowledge Work
New Tasks: More Realistic Knowledge Work|News screenshot

The update adds two major evaluations focused on more complex and realistic knowledge work and long-context document reasoning:

  1. AA-Briefcase: An in-house Artificial Analysis evaluation with a private held-out test set. It tests models on realistic agentic knowledge work tasks in complex projects built by industry experts. Models are evaluated on multi-week knowledge work projects, each with many linked tasks and thousands of input source files. Scoring combines rubric-based and pairwise grading across verifiable task success, analytical quality, and presentation quality.

  2. GDP.pdf: Created by Surge AI, GDP.pdf evaluates single-turn professional document reasoning across 100 PDFs, 10 domains, and 4,592 pages. Models must synthesize evidence from text, tables, charts, footnotes, and exclusions. Responses are graded against 1,275 expert-authored atomic criteria; the headline All-pass Rate credits a task only when every criterion is satisfied.

One notable change is the removal of GPQA Diamond. Artificial Analysis describes it as an exceptional scientific reasoning evaluation that has now been saturated. That highlights a broader benchmark problem: once frontier models approach ceiling performance on a task, evaluation providers need harder, more realistic, and less gameable tests.

Grading Infrastructure: Stability and Robustness

v4.2 also upgrades the grading pipeline to improve scoring accuracy and stability:

  • AA-LCR v1.1: Adds a grading system prompt and corrects errors and ambiguities in answer keys.
  • GDPval-AA v2 and AA-Briefcase: Improve sampling and re-anchor the Elo scale, making ratings more stable as new models are added.
  • SciCode: Improves grading sandbox robustness so slow but correct code is not counted as a failure.

On anti-gaming, private held-out data—including AA-Briefcase, AA-Omniscience, and solutions for CritPt—now accounts for 40% of the Index weighting. For labs that optimize heavily against published leaderboards, this reduces the value of training or tuning directly against visible test items.

Model Performance and Efficiency Frontiers

Model Performance and Efficiency Frontiers
Model Performance and Efficiency Frontiers|News screenshot

The updated benchmark results include several headline findings:

  • Overall ranking: Anthropic’s Claude Fable 5.1 leads the Index, followed by OpenAI’s GPT-6 Astra. GPT-6 Astra shows a 4-point gain over GPT-5.6 Sol on the Intelligence Index. Meta is the third-ranked lab, followed by SpaceXAI, Moonshot/Kimi, Z.AI, and Google.

  • Cost per Task frontier: The updated Pareto frontier is shared by four labs: Anthropic, OpenAI, Meta, and Z.AI.

  • Output token frontier: Among models scoring at least 25 on the Index, GPT-6 Astra is more token-efficient than almost every other model near the intelligence frontier, with Claude Fable 5.1, Grok 4.5, and Gemini 3.5 Flash-Lite at either end of the curve.

  • Task-specific results:

    • AA-Briefcase: Claude Fable 5.1 and Opus 5 lead, followed by GPT-6 Astra and Muse Spark 1.3. GPT-6 Astra shows a gain of about 85 Elo points over GPT-5.6 Sol on this evaluation.
    • GDP.pdf: OpenAI leads with GPT-6 Astra at 33.2% All-pass Rate, followed by GPT-5.6 Sol at 28.2% and Claude Fable 5.1 at 26.2%.

The results underline a familiar trade-off: overall intelligence, token efficiency, and cost per task do not always point to the same model. GPT-6 Astra stands out on output token efficiency, while the cost frontier remains shared across multiple labs. Practical model selection still depends on workload, pricing, latency, and reliability requirements.

Reader Recommendations

Reader Recommendations
Reader Recommendations|News screenshot

  • For practitioners: If your workload involves complex knowledge work, multi-document synthesis, professional analysis, or agentic task execution, AA-Briefcase and GDP.pdf are likely more relevant than older short-form benchmarks.

  • For benchmark users: A 40% private held-out weighting makes the Index harder to game and potentially more predictive of real-world performance, but it also reduces external visibility into exact test items. Teams should treat the leaderboard as one signal and validate models on their own task suites before deployment.

Final Thought

Intelligence Index v4.2 reflects a broader shift in AI evaluation: from public academic-style test sets toward industrial benchmarks designed around realism and anti-gaming. The important change is not just the addition of two evaluations, but the higher weight given to private held-out tests, long-context document reasoning, and agentic knowledge work. As v5 approaches, these dimensions are likely to become even more central to frontier model assessment.