Releases
123 matching — model launches, API releases, product updates, financial events and news.
September 20265
Sep 30, 2026
Cohere introduces RCP-nDCG@10 retrieval metricResearch paper
Cohere proposed RCP-nDCG@10 as a retrieval-quality measure designed to align more closely with human relevance judgments.
Sep 23, 2026
Anthropic opens life-sciences lab and reports Claude-assisted enzyme discoveryResearch paper
Anthropic introduced a new life-sciences research group and laboratory and reported that Claude agents identified a novel array-associated reverse-transcriptase system with CRISPR-like repeats.
Sep 14, 2026
Modal scales to one million concurrent SandboxesBenchmark result
Modal reported support for up to one million concurrent Sandboxes.
Sep 10, 2026
Anthropic publishes September 2026 threat intelligence reportSafety report
Anthropic published case studies on misuse of Claude across cyber, influence, surveillance, fraud, biological, weapons, and model-distillation activity detected from December 2025 through August 2026.
Sep 1, 2026
Independent biosecurity evaluation published for Grok 4.6Safety report
SpaceXAI published LatchBio evaluation results reporting strong detection and refusal performance on adversarial biological tasks.
June 20262
Jun 22, 2026
CoreWeave sets MLPerf Training v6.0 recordsBenchmark result
CoreWeave reported record MLPerf Training v6.0 results, including training DeepSeek-V3 in approximately two minutes on a large NVIDIA GB300 NVL72 cluster.
Jun 4, 2026
Project Granite Switch releasedResearch paper
IBM introduced a system for composing model adapters into deployable specialized models.
May 20262
May 13, 2026
Computer security architecture publishedSafety report
Perplexity published details of the security architecture and safeguards built into Computer.
May 11, 2026
CoreWeave ranks first for Kimi K2.6 inference speed and price-performanceBenchmark result
Artificial Analysis data placed CoreWeave first for Kimi K2.6 inference throughput and price-performance among evaluated providers.
April 20261
Apr 22, 2026
Frontier AI accuracy research publishedResearch paper
Perplexity published its approach to building and evaluating accuracy in frontier AI systems.
March 20264
Mar 9, 2026
Agentic Rubrics introducedResearch paper
Scale Labs introduced repository-grounded rubrics for evaluating and reranking candidate software patches without executing tests.
Mar 9, 2026
VeRO evaluation framework introducedResearch paper
Scale Labs introduced VeRO, a reproducible harness using versioned agent snapshots, controlled budgets, structured traces, and reference procedures.
Mar 4, 2026
SWE Atlas launches with Codebase QnABenchmark result
Scale launched SWE Atlas, initially releasing Codebase QnA to evaluate how coding agents investigate and reason about real software systems.
Mar 2026
MoBA research implementation releasedResearch paper
Moonshot published Mixture-of-Block Attention code and research for efficient long-context modeling.
January 20261
Jan 23, 2026
Long-Horizon Augmented Workflows releasedBenchmark result
Scale researchers released LHAW, a framework for generating underspecified long-horizon tasks and evaluating whether agents clarify ambiguity and recover performance.
December 20252
Dec 31, 2025
xAI publishes Frontier Artificial Intelligence FrameworkSafety report
xAI published its frontier-AI risk-management framework covering capability evaluation, safeguards and model development practices.
Dec 2, 2025
BrowseSafe browser security research releasedResearch paper
Perplexity released BrowseSafe research on detecting and mitigating threats in AI browsers.
October 20252
Oct 20, 2025
DeepSeek OCR research family introducedResearch paper
DeepSeek introduced its open OCR and document-understanding research family based on visual-text compression.
Oct 2025
Kimi Linear architecture introducedResearch paper
Moonshot introduced Kimi Linear and Kimi Delta Attention as a hybrid linear-attention approach.
September 20252
Sep 19, 2025
MCP Atlas releasedBenchmark result
Scale released MCP Atlas to evaluate how well AI agents combine tools across real Model Context Protocol servers.
Sep 19, 2025
SWE-Bench Pro releasedBenchmark result
Scale released a harder, contamination-resistant benchmark using real and commercial repositories to measure software agents on complex multi-file engineering tasks.
August 20252
Aug 20, 2025
Grok 4 model card publishedSafety report
xAI published the Grok 4 model card covering evaluations, safeguards, abuse potential and dual-use capabilities.
Aug 5, 2025
Genie 3 world model introducedResearch paper
Google DeepMind introduces Genie 3, a general-purpose world model that generates interactive environments in real time.
July 20251
Jul 9, 2025
xAI removes antisemitic Grok posts and updates safeguardsSafety report
xAI removed inappropriate Grok posts containing antisemitic tropes and praise for Hitler, tightened hate-speech controls and later explained that a system update had made the bot too susceptible to extremist user content.
June 20251
Jun 10, 2025
Sonar inference acceleratedResearch paper
Perplexity published work on accelerating Sonar inference through speculative decoding.

