Releases
11 matching — model launches, API releases, product updates, financial events and news.
March 20263
Mar 9, 2026
Agentic Rubrics introducedResearch paper
Scale Labs introduced repository-grounded rubrics for evaluating and reranking candidate software patches without executing tests.
Mar 9, 2026
VeRO evaluation framework introducedResearch paper
Scale Labs introduced VeRO, a reproducible harness using versioned agent snapshots, controlled budgets, structured traces, and reference procedures.
Mar 4, 2026
SWE Atlas launches with Codebase QnABenchmark result
Scale launched SWE Atlas, initially releasing Codebase QnA to evaluate how coding agents investigate and reason about real software systems.
January 20261
Jan 23, 2026
Long-Horizon Augmented Workflows releasedBenchmark result
Scale researchers released LHAW, a framework for generating underspecified long-horizon tasks and evaluating whether agents clarify ambiguity and recover performance.
September 20252
Sep 19, 2025
MCP Atlas releasedBenchmark result
Scale released MCP Atlas to evaluate how well AI agents combine tools across real Model Context Protocol servers.
Sep 19, 2025
SWE-Bench Pro releasedBenchmark result
Scale released a harder, contamination-resistant benchmark using real and commercial repositories to measure software agents on complex multi-file engineering tasks.
February 20252
Feb 11, 2025
MASK belief-alignment benchmark releasedBenchmark result
Scale published MASK as part of its safety research program for testing whether models knowingly state beliefs inconsistent with their internal representations.
Feb 11, 2025
FORTRESS benchmark releasedBenchmark result
Scale released FORTRESS to evaluate frontier-model safeguards against dual-use national-security and public-safety risks.
January 20251
Jan 23, 2025
Humanity's Last Exam results and benchmark releasedBenchmark result
Scale and the Center for AI Safety published Humanity’s Last Exam, assembled from nearly 1,000 contributors across more than 500 institutions in 50 countries.
September 20241
Sep 2024
EnigmaEval benchmark releasedBenchmark result
Scale and research collaborators introduced EnigmaEval, a multimodal reasoning benchmark built from novel puzzle-competition problems.
August 20241
Aug 2024
MultiChallenge benchmark releasedBenchmark result
Scale released MultiChallenge to measure model performance on realistic multi-turn conversation problems involving memory, instruction retention, editing, and self-consistency.

