Skip to content
AIMarketCap

Releases

123 matching — model launches, API releases, product updates, financial events and news.

Clear filters

September 20265

Sep 30, 2026Cohere introduces RCP-nDCG@10 retrieval metricCohere

Cohere proposed RCP-nDCG@10 as a retrieval-quality measure designed to align more closely with human relevance judgments.

Cohere Research paperCohere
Sep 23, 2026Anthropic opens life-sciences lab and reports Claude-assisted enzyme discoveryAnthropic

Anthropic introduced a new life-sciences research group and laboratory and reported that Claude agents identified a novel array-associated reverse-transcriptase system with CRISPR-like repeats.

Anthropic Research paperAnthropic
Sep 14, 2026Modal scales to one million concurrent SandboxesModal Labs · Modal Sandboxes

Modal reported support for up to one million concurrent Sandboxes.

Sep 10, 2026Anthropic publishes September 2026 threat intelligence reportAnthropic · September 2026

Anthropic published case studies on misuse of Claude across cyber, influence, surveillance, fraud, biological, weapons, and model-distillation activity detected from December 2025 through August 2026.

Anthropic Safety reportAnthropic
Sep 1, 2026Independent biosecurity evaluation published for Grok 4.6SpaceXAI · Grok 4.6 · 4.6

SpaceXAI published LatchBio evaluation results reporting strong detection and refusal performance on adversarial biological tasks.

June 20262

Jun 22, 2026CoreWeave sets MLPerf Training v6.0 recordsCoreWeave · CoreWeave Compute

CoreWeave reported record MLPerf Training v6.0 results, including training DeepSeek-V3 in approximately two minutes on a large NVIDIA GB300 NVL72 cluster.

Jun 4, 2026Project Granite Switch releasedIBM · Project Granite Switch

IBM introduced a system for composing model adapters into deployable specialized models.

May 20262

May 13, 2026Computer security architecture publishedPerplexity · Perplexity Computer

Perplexity published details of the security architecture and safeguards built into Computer.

May 11, 2026CoreWeave ranks first for Kimi K2.6 inference speed and price-performanceCoreWeave · CoreWeave Inference

Artificial Analysis data placed CoreWeave first for Kimi K2.6 inference throughput and price-performance among evaluated providers.

April 20261

Apr 22, 2026Frontier AI accuracy research publishedPerplexity · Perplexity Answer Engine

Perplexity published its approach to building and evaluating accuracy in frontier AI systems.

March 20264

Mar 9, 2026Agentic Rubrics introducedScale AI · Agentic Rubrics

Scale Labs introduced repository-grounded rubrics for evaluating and reranking candidate software patches without executing tests.

Mar 9, 2026VeRO evaluation framework introducedScale AI · VeRO

Scale Labs introduced VeRO, a reproducible harness using versioned agent snapshots, controlled budgets, structured traces, and reference procedures.

Scale AI Research paperScale AIVeRO
Mar 4, 2026SWE Atlas launches with Codebase QnAScale AI · SWE Atlas

Scale launched SWE Atlas, initially releasing Codebase QnA to evaluate how coding agents investigate and reason about real software systems.

Mar 2026MoBA research implementation releasedMoonshot AI · MoBA

Moonshot published Mixture-of-Block Attention code and research for efficient long-context modeling.

January 20261

Jan 23, 2026Long-Horizon Augmented Workflows releasedScale AI · Long-Horizon Augmented Workflows

Scale researchers released LHAW, a framework for generating underspecified long-horizon tasks and evaluating whether agents clarify ambiguity and recover performance.

December 20252

Dec 31, 2025xAI publishes Frontier Artificial Intelligence FrameworkSpaceXAI

xAI published its frontier-AI risk-management framework covering capability evaluation, safeguards and model development practices.

SpaceXAI Safety reportSpaceXAI
Dec 2, 2025BrowseSafe browser security research releasedPerplexity · Comet

Perplexity released BrowseSafe research on detecting and mitigating threats in AI browsers.

October 20252

Oct 20, 2025DeepSeek OCR research family introducedDeepSeek · DeepSeek OCR

DeepSeek introduced its open OCR and document-understanding research family based on visual-text compression.

Oct 2025Kimi Linear architecture introducedMoonshot AI · Kimi Linear

Moonshot introduced Kimi Linear and Kimi Delta Attention as a hybrid linear-attention approach.

September 20252

Sep 19, 2025MCP Atlas releasedScale AI · MCP Atlas

Scale released MCP Atlas to evaluate how well AI agents combine tools across real Model Context Protocol servers.

Sep 19, 2025SWE-Bench Pro releasedScale AI · SWE-Bench Pro

Scale released a harder, contamination-resistant benchmark using real and commercial repositories to measure software agents on complex multi-file engineering tasks.

August 20252

Aug 20, 2025Grok 4 model card publishedSpaceXAI · Grok 4 · 4

xAI published the Grok 4 model card covering evaluations, safeguards, abuse potential and dual-use capabilities.

Aug 5, 2025Genie 3 world model introducedGoogle (Alphabet) · Genie 3

Google DeepMind introduces Genie 3, a general-purpose world model that generates interactive environments in real time.

July 20251

Jul 9, 2025xAI removes antisemitic Grok posts and updates safeguardsSpaceXAI · Grok

xAI removed inappropriate Grok posts containing antisemitic tropes and praise for Hitler, tightened hate-speech controls and later explained that a system update had made the bot too susceptible to extremist user content.

June 20252

Jun 10, 2025Sonar inference acceleratedPerplexity · Sonar

Perplexity published work on accelerating Sonar inference through speculative decoding.

Jun 4, 2025CoreWeave, NVIDIA, and IBM set Blackwell MLPerf training recordsCoreWeave · CoreWeave Compute

The companies submitted the largest NVIDIA Blackwell MLPerf Training v5.0 cluster and reported record Llama 3.1 405B training performance.

May 20252

May 16, 2025Unauthorized Grok prompt modification disclosedSpaceXAI · Grok

xAI disclosed that an unauthorized system-prompt change caused Grok to inject unrelated claims about South African racial politics; it published prompts and announced stronger review and monitoring controls.

May 14, 2025AlphaEvolve introducedGoogle (Alphabet) · AlphaEvolve

Google DeepMind introduces AlphaEvolve, a Gemini-powered evolutionary coding agent for algorithm discovery and optimization.

April 20251

Apr 2, 2025CoreWeave posts first cloud NVIDIA GB200 MLPerf inference resultsCoreWeave · CoreWeave Compute

CoreWeave submitted the first cloud-provider MLPerf Inference v5.0 results for NVIDIA GB200, reporting more than 800 tokens per second on Llama 3.1 405B.

March 20252

Mar 27, 2025Tracing the thoughts of a large language model publishedAnthropic · Claude

Anthropic published interpretability research tracing internal mechanisms underlying multilingual reasoning, planning, and unfaithful explanations.

Mar 2025Kimina-Prover project introducedMoonshot AI · Kimina-Prover

Moonshot introduced the Kimina-Prover line of open models for formal theorem proving.

February 20257

Feb 19, 2025Muse world and action model introducedMicrosoft · Muse

Microsoft Research and Xbox introduce Muse, a generative world and action model trained on gameplay sequences.

Feb 11, 2025MASK belief-alignment benchmark releasedScale AI · MASK

Scale published MASK as part of its safety research program for testing whether models knowingly state beliefs inconsistent with their internal representations.

Scale AI Benchmark resultScale AIMASK
Feb 11, 2025FORTRESS benchmark releasedScale AI · FORTRESS

Scale released FORTRESS to evaluate frontier-model safeguards against dual-use national-security and public-safety risks.

Feb 10, 2025Anthropic Economic Index launchedAnthropic

Anthropic launched the Economic Index to measure how AI is being used across occupations and tasks.

Anthropic Research paperAnthropic
Feb 7, 2025PARTNR benchmark and dataset releasedMeta · PARTNR

Meta released PARTNR, a benchmark, dataset and planning model for human-robot collaboration on household tasks.

Meta AI Research paperMetaPARTNR
Feb 6, 2025Brain2Qwerty research publishedMeta · Brain2Qwerty · v1

Meta researchers introduced a non-invasive deep-learning system for decoding typed sentences from EEG and MEG brain recordings.

Feb 3, 2025Constitutional Classifiers research publishedAnthropic · Claude

Anthropic published and tested Constitutional Classifiers, a safeguard approach designed to resist broad classes of jailbreaks.

January 20255

Jan 23, 2025Humanity's Last Exam results and benchmark releasedScale AI · Humanity's Last Exam

Scale and the Center for AI Safety published Humanity’s Last Exam, assembled from nearly 1,000 contributors across more than 500 institutions in 50 countries.

Jan 20, 2025DeepSeek R1 reasoning family launchesDeepSeek · DeepSeek R1

DeepSeek introduced its reinforcement-learning reasoning model family and technical report.

Jan 16, 2025MatterGen research publishedMicrosoft · MatterGen

Microsoft Research publishes MatterGen, a generative model that designs stable inorganic materials conditioned on target properties.

Jan 2025Kimi k1 reasoning model disclosedMoonshot AI · Kimi k1

Moonshot described k1 as a long-context reinforcement-learning reasoning system.

2025TRIBE research model recognized at Algonauts 2025Meta · TRIBE · v1

Meta's first TRIBE brain-response model formed the foundation for its award-winning Algonauts 2025 entry; Meta's later TRIBE v2 announcement identifies the original model as 2025 work.

Meta AI Benchmark resultMetaTRIBE

December 20244

Dec 26, 2024DeepSeek V3 family launchesDeepSeek · DeepSeek V3

DeepSeek introduced its 671B-parameter V3 mixture-of-experts architecture.

Dec 18, 2024Alignment faking research publishedAnthropic · Claude 3.5 Sonnet (October 2024)

Anthropic and Redwood Research published evidence of alignment-faking behavior in a controlled Claude 3 Opus experiment.

Dec 4, 2024BioEmu generative protein model introducedMicrosoft · BioEmu

Microsoft Research introduces BioEmu for efficiently sampling diverse protein conformational ensembles.

Dec 4, 2024Genie 2 world model introducedGoogle (Alphabet) · Genie 2

Google DeepMind introduces Genie 2, capable of generating diverse playable 3D environments from a prompt image.

November 20241

Nov 2024Kimi k0-math research model disclosedMoonshot AI · Kimi k0-math

Moonshot disclosed k0-math as an experimental reinforcement-learning model for mathematical reasoning.

October 20242

Oct 18, 2024Janus multimodal family introducedDeepSeek · Janus

DeepSeek introduced the Janus family for unified multimodal understanding and generation.

Oct 18, 2024Spirit LM introducedMeta · Spirit LM

Meta released research on a multimodal language model that freely mixes text and speech.

Meta AI Research paperMetaSpirit LM
1–50 of 123 releases