Skip to content
AIMarketCap

Releases

135 matching — model launches, API releases, product updates, financial events and news.

Clear filters

October 20262

Oct 5, 2026AWS Continuum achieves 89% end-to-end success on CyberGym-E2EAmazon · AWS Continuum

AWS reported that Continuum for code vulnerabilities passed 819 of 920 CyberGym-E2E tasks within the benchmark's 90-minute limit, achieving an 89.0% end-to-end success rate, 23.1 percentage points above the previous public high of 65.9%.

Oct 1, 2026Connector library architecture publishedHarvey · Harvey Integrations

Harvey published its architecture for connecting agents to external tools and systems.

September 20266

Sep 30, 2026Cohere introduces RCP-nDCG@10 retrieval metricCohere

Cohere proposed RCP-nDCG@10 as a retrieval-quality measure designed to align more closely with human relevance judgments.

Cohere Research paperCohere
Sep 23, 2026Anthropic opens life-sciences lab and reports Claude-assisted enzyme discoveryAnthropic

Anthropic introduced a new life-sciences research group and laboratory and reported that Claude agents identified a novel array-associated reverse-transcriptase system with CRISPR-like repeats.

Anthropic Research paperAnthropic
Sep 17, 2026H3 Max ranks first in fal evaluationsFal AI · H3 Max

fal reported H3 Max ranked first in its human-preference evaluation across quality, prompt understanding, and aesthetics, with top results also cited from independent benchmarks.

fal Benchmark resultFal AIH3 Max
Sep 14, 2026Modal scales to one million concurrent SandboxesModal Labs · Modal Sandboxes

Modal reported support for up to one million concurrent Sandboxes.

Sep 10, 2026Anthropic publishes September 2026 threat intelligence reportAnthropic · September 2026

Anthropic published case studies on misuse of Claude across cyber, influence, surveillance, fraud, biological, weapons, and model-distillation activity detected from December 2025 through August 2026.

Anthropic Safety reportAnthropic
Sep 1, 2026Independent biosecurity evaluation published for Grok 4.6SpaceXAI · Grok 4.6 · 4.6

SpaceXAI published LatchBio evaluation results reporting strong detection and refusal performance on adversarial biological tasks.

July 20261

Jul 20, 2026Cursor publishes agent-swarm economics researchCursor (Anysphere) · Cursor Agent Swarms

Cursor published research on the economics and scaling behavior of large coding-agent swarms.

June 20265

Jun 22, 2026CoreWeave sets MLPerf Training v6.0 recordsCoreWeave · CoreWeave Compute

CoreWeave reported record MLPerf Training v6.0 results, including training DeepSeek-V3 in approximately two minutes on a large NVIDIA GB300 NVL72 cluster.

Jun 4, 2026Project Granite Switch releasedIBM · Project Granite Switch

IBM introduced a system for composing model adapters into deployable specialized models.

Jun 2, 2026Cursor publishes cloud-agent operating lessonsCursor (Anysphere) · Cursor Cloud Agents

Cursor published lessons from operating cloud coding agents at scale.

Jun 1, 2026Harvey LAB legal-agent benchmark introducedHarvey · Harvey LAB

Harvey introduced LAB, its benchmark for comparing legal-agent quality, practice-area performance, and cost efficiency.

Harvey Benchmark resultHarveyHarvey LAB
Jun 1, 2026Cloud agent infrastructure research publishedHarvey · Harvey Agents

Harvey described the infrastructure built to execute secure legal agents at scale.

May 20262

May 13, 2026Computer security architecture publishedPerplexity · Perplexity Computer

Perplexity published details of the security architecture and safeguards built into Computer.

May 11, 2026CoreWeave ranks first for Kimi K2.6 inference speed and price-performanceCoreWeave · CoreWeave Inference

Artificial Analysis data placed CoreWeave first for Kimi K2.6 inference throughput and price-performance among evaluated providers.

April 20262

Apr 30, 2026Cursor publishes agent-harness researchCursor (Anysphere) · Cursor Agent

Cursor described its continuously improving agent harness and the feedback loops used to develop coding agents.

Apr 22, 2026Frontier AI accuracy research publishedPerplexity · Perplexity Answer Engine

Perplexity published its approach to building and evaluating accuracy in frontier AI systems.

March 20264

Mar 9, 2026Agentic Rubrics introducedScale AI · Agentic Rubrics

Scale Labs introduced repository-grounded rubrics for evaluating and reranking candidate software patches without executing tests.

Mar 9, 2026VeRO evaluation framework introducedScale AI · VeRO

Scale Labs introduced VeRO, a reproducible harness using versioned agent snapshots, controlled budgets, structured traces, and reference procedures.

Scale AI Research paperScale AIVeRO
Mar 4, 2026SWE Atlas launches with Codebase QnAScale AI · SWE Atlas

Scale launched SWE Atlas, initially releasing Codebase QnA to evaluate how coding agents investigate and reason about real software systems.

Mar 2026MoBA research implementation releasedMoonshot AI · MoBA

Moonshot published Mixture-of-Block Attention code and research for efficient long-context modeling.

January 20262

Jan 23, 2026Long-Horizon Augmented Workflows releasedScale AI · Long-Horizon Augmented Workflows

Scale researchers released LHAW, a framework for generating underspecified long-horizon tasks and evaluating whether agents clarify ambiguity and recover performance.

Jan 21, 2026Mercor releases APEX-AgentsMercor · APEX-Agents · 1.0

Mercor released the open APEX-Agents benchmark for long-horizon professional work; frontier agents completed fewer than 25% of tasks on the initial evaluation.

Mercor Benchmark resultMercorAPEX-Agents

December 20251

Dec 31, 2025xAI publishes Frontier Artificial Intelligence FrameworkSpaceXAI

xAI published its frontier-AI risk-management framework covering capability evaluation, safeguards and model development practices.

SpaceXAI Safety reportSpaceXAI
1–25 of 135 releases