Skip to content
AIMarketCap

Releases

11 matching — model launches, API releases, product updates, financial events and news.

Clear filters

March 20263

Mar 9, 2026Agentic Rubrics introducedScale AI · Agentic Rubrics

Scale Labs introduced repository-grounded rubrics for evaluating and reranking candidate software patches without executing tests.

Mar 9, 2026VeRO evaluation framework introducedScale AI · VeRO

Scale Labs introduced VeRO, a reproducible harness using versioned agent snapshots, controlled budgets, structured traces, and reference procedures.

Scale AI Research paperScale AIVeRO
Mar 4, 2026SWE Atlas launches with Codebase QnAScale AI · SWE Atlas

Scale launched SWE Atlas, initially releasing Codebase QnA to evaluate how coding agents investigate and reason about real software systems.

January 20261

Jan 23, 2026Long-Horizon Augmented Workflows releasedScale AI · Long-Horizon Augmented Workflows

Scale researchers released LHAW, a framework for generating underspecified long-horizon tasks and evaluating whether agents clarify ambiguity and recover performance.

September 20252

Sep 19, 2025MCP Atlas releasedScale AI · MCP Atlas

Scale released MCP Atlas to evaluate how well AI agents combine tools across real Model Context Protocol servers.

Sep 19, 2025SWE-Bench Pro releasedScale AI · SWE-Bench Pro

Scale released a harder, contamination-resistant benchmark using real and commercial repositories to measure software agents on complex multi-file engineering tasks.

February 20252

Feb 11, 2025MASK belief-alignment benchmark releasedScale AI · MASK

Scale published MASK as part of its safety research program for testing whether models knowingly state beliefs inconsistent with their internal representations.

Scale AI Benchmark resultScale AIMASK
Feb 11, 2025FORTRESS benchmark releasedScale AI · FORTRESS

Scale released FORTRESS to evaluate frontier-model safeguards against dual-use national-security and public-safety risks.

January 20251

Jan 23, 2025Humanity's Last Exam results and benchmark releasedScale AI · Humanity's Last Exam

Scale and the Center for AI Safety published Humanity’s Last Exam, assembled from nearly 1,000 contributors across more than 500 institutions in 50 countries.

September 20241

Sep 2024EnigmaEval benchmark releasedScale AI · EnigmaEval

Scale and research collaborators introduced EnigmaEval, a multimodal reasoning benchmark built from novel puzzle-competition problems.

August 20241

Aug 2024MultiChallenge benchmark releasedScale AI · MultiChallenge

Scale released MultiChallenge to measure model performance on realistic multi-turn conversation problems involving memory, instruction retention, editing, and self-consistency.

1–11 of 11 releases
AI releases — calendar and timeline of models, APIs, funding and news · The Artificial Index