Releases
129 matching — model launches, API releases, product updates, financial events and news.
October 20261
Oct 1, 2026Connector library architecture publishedResearch paper
Harvey published its architecture for connecting agents to external tools and systems.
September 20265
Sep 30, 2026
Cohere introduces RCP-nDCG@10 retrieval metricResearch paper
Cohere proposed RCP-nDCG@10 as a retrieval-quality measure designed to align more closely with human relevance judgments.
Sep 23, 2026
Anthropic opens life-sciences lab and reports Claude-assisted enzyme discoveryResearch paper
Anthropic introduced a new life-sciences research group and laboratory and reported that Claude agents identified a novel array-associated reverse-transcriptase system with CRISPR-like repeats.
Sep 14, 2026
Modal scales to one million concurrent SandboxesBenchmark result
Modal reported support for up to one million concurrent Sandboxes.
Sep 10, 2026
Anthropic publishes September 2026 threat intelligence reportSafety report
Anthropic published case studies on misuse of Claude across cyber, influence, surveillance, fraud, biological, weapons, and model-distillation activity detected from December 2025 through August 2026.
Sep 1, 2026
Independent biosecurity evaluation published for Grok 4.6Safety report
SpaceXAI published LatchBio evaluation results reporting strong detection and refusal performance on adversarial biological tasks.
July 20261
Jul 20, 2026Cursor publishes agent-swarm economics researchResearch paper
Cursor published research on the economics and scaling behavior of large coding-agent swarms.
June 20265
Jun 22, 2026
CoreWeave sets MLPerf Training v6.0 recordsBenchmark result
CoreWeave reported record MLPerf Training v6.0 results, including training DeepSeek-V3 in approximately two minutes on a large NVIDIA GB300 NVL72 cluster.
Jun 4, 2026
Project Granite Switch releasedResearch paper
IBM introduced a system for composing model adapters into deployable specialized models.
Jun 2, 2026Cursor publishes cloud-agent operating lessonsResearch paper
Cursor published lessons from operating cloud coding agents at scale.
Jun 1, 2026Harvey LAB legal-agent benchmark introducedBenchmark result
Harvey introduced LAB, its benchmark for comparing legal-agent quality, practice-area performance, and cost efficiency.
Jun 1, 2026Cloud agent infrastructure research publishedResearch paper
Harvey described the infrastructure built to execute secure legal agents at scale.
May 20262
May 13, 2026
Computer security architecture publishedSafety report
Perplexity published details of the security architecture and safeguards built into Computer.
May 11, 2026
CoreWeave ranks first for Kimi K2.6 inference speed and price-performanceBenchmark result
Artificial Analysis data placed CoreWeave first for Kimi K2.6 inference throughput and price-performance among evaluated providers.
April 20262
Apr 30, 2026Cursor publishes agent-harness researchResearch paper
Cursor described its continuously improving agent harness and the feedback loops used to develop coding agents.
Apr 22, 2026
Frontier AI accuracy research publishedResearch paper
Perplexity published its approach to building and evaluating accuracy in frontier AI systems.
March 20264
Mar 9, 2026
Agentic Rubrics introducedResearch paper
Scale Labs introduced repository-grounded rubrics for evaluating and reranking candidate software patches without executing tests.
Mar 9, 2026
VeRO evaluation framework introducedResearch paper
Scale Labs introduced VeRO, a reproducible harness using versioned agent snapshots, controlled budgets, structured traces, and reference procedures.
Mar 4, 2026
SWE Atlas launches with Codebase QnABenchmark result
Scale launched SWE Atlas, initially releasing Codebase QnA to evaluate how coding agents investigate and reason about real software systems.
Mar 2026
MoBA research implementation releasedResearch paper
Moonshot published Mixture-of-Block Attention code and research for efficient long-context modeling.
January 20261
Jan 23, 2026
Long-Horizon Augmented Workflows releasedBenchmark result
Scale researchers released LHAW, a framework for generating underspecified long-horizon tasks and evaluating whether agents clarify ambiguity and recover performance.
December 20252
Dec 31, 2025
xAI publishes Frontier Artificial Intelligence FrameworkSafety report
xAI published its frontier-AI risk-management framework covering capability evaluation, safeguards and model development practices.
Dec 2, 2025
BrowseSafe browser security research releasedResearch paper
Perplexity released BrowseSafe research on detecting and mitigating threats in AI browsers.
October 20252
Oct 20, 2025
DeepSeek OCR research family introducedResearch paper
DeepSeek introduced its open OCR and document-understanding research family based on visual-text compression.
Oct 2025
Kimi Linear architecture introducedResearch paper
Moonshot introduced Kimi Linear and Kimi Delta Attention as a hybrid linear-attention approach.

