Releases
123 matching — model launches, API releases, product updates, financial events and news.
September 20265
Sep 30, 2026
Cohere introduces RCP-nDCG@10 retrieval metricResearch paper
Cohere proposed RCP-nDCG@10 as a retrieval-quality measure designed to align more closely with human relevance judgments.
Sep 23, 2026
Anthropic opens life-sciences lab and reports Claude-assisted enzyme discoveryResearch paper
Anthropic introduced a new life-sciences research group and laboratory and reported that Claude agents identified a novel array-associated reverse-transcriptase system with CRISPR-like repeats.
Sep 14, 2026
Modal scales to one million concurrent SandboxesBenchmark result
Modal reported support for up to one million concurrent Sandboxes.
Sep 10, 2026
Anthropic publishes September 2026 threat intelligence reportSafety report
Anthropic published case studies on misuse of Claude across cyber, influence, surveillance, fraud, biological, weapons, and model-distillation activity detected from December 2025 through August 2026.
Sep 1, 2026
Independent biosecurity evaluation published for Grok 4.6Safety report
SpaceXAI published LatchBio evaluation results reporting strong detection and refusal performance on adversarial biological tasks.
June 20262
Jun 22, 2026
CoreWeave sets MLPerf Training v6.0 recordsBenchmark result
CoreWeave reported record MLPerf Training v6.0 results, including training DeepSeek-V3 in approximately two minutes on a large NVIDIA GB300 NVL72 cluster.
Jun 4, 2026
Project Granite Switch releasedResearch paper
IBM introduced a system for composing model adapters into deployable specialized models.
May 20262
May 13, 2026
Computer security architecture publishedSafety report
Perplexity published details of the security architecture and safeguards built into Computer.
May 11, 2026
CoreWeave ranks first for Kimi K2.6 inference speed and price-performanceBenchmark result
Artificial Analysis data placed CoreWeave first for Kimi K2.6 inference throughput and price-performance among evaluated providers.
April 20261
Apr 22, 2026
Frontier AI accuracy research publishedResearch paper
Perplexity published its approach to building and evaluating accuracy in frontier AI systems.
March 20264
Mar 9, 2026
Agentic Rubrics introducedResearch paper
Scale Labs introduced repository-grounded rubrics for evaluating and reranking candidate software patches without executing tests.
Mar 9, 2026
VeRO evaluation framework introducedResearch paper
Scale Labs introduced VeRO, a reproducible harness using versioned agent snapshots, controlled budgets, structured traces, and reference procedures.
Mar 4, 2026
SWE Atlas launches with Codebase QnABenchmark result
Scale launched SWE Atlas, initially releasing Codebase QnA to evaluate how coding agents investigate and reason about real software systems.
Mar 2026
MoBA research implementation releasedResearch paper
Moonshot published Mixture-of-Block Attention code and research for efficient long-context modeling.
January 20261
Jan 23, 2026
Long-Horizon Augmented Workflows releasedBenchmark result
Scale researchers released LHAW, a framework for generating underspecified long-horizon tasks and evaluating whether agents clarify ambiguity and recover performance.
December 20252
Dec 31, 2025
xAI publishes Frontier Artificial Intelligence FrameworkSafety report
xAI published its frontier-AI risk-management framework covering capability evaluation, safeguards and model development practices.
Dec 2, 2025
BrowseSafe browser security research releasedResearch paper
Perplexity released BrowseSafe research on detecting and mitigating threats in AI browsers.
October 20252
Oct 20, 2025
DeepSeek OCR research family introducedResearch paper
DeepSeek introduced its open OCR and document-understanding research family based on visual-text compression.
Oct 2025
Kimi Linear architecture introducedResearch paper
Moonshot introduced Kimi Linear and Kimi Delta Attention as a hybrid linear-attention approach.
September 20252
Sep 19, 2025
MCP Atlas releasedBenchmark result
Scale released MCP Atlas to evaluate how well AI agents combine tools across real Model Context Protocol servers.
Sep 19, 2025
SWE-Bench Pro releasedBenchmark result
Scale released a harder, contamination-resistant benchmark using real and commercial repositories to measure software agents on complex multi-file engineering tasks.
August 20252
Aug 20, 2025
Grok 4 model card publishedSafety report
xAI published the Grok 4 model card covering evaluations, safeguards, abuse potential and dual-use capabilities.
Aug 5, 2025
Genie 3 world model introducedResearch paper
Google DeepMind introduces Genie 3, a general-purpose world model that generates interactive environments in real time.
July 20251
Jul 9, 2025
xAI removes antisemitic Grok posts and updates safeguardsSafety report
xAI removed inappropriate Grok posts containing antisemitic tropes and praise for Hitler, tightened hate-speech controls and later explained that a system update had made the bot too susceptible to extremist user content.
June 20252
Jun 10, 2025
Sonar inference acceleratedResearch paper
Perplexity published work on accelerating Sonar inference through speculative decoding.
Jun 4, 2025
CoreWeave, NVIDIA, and IBM set Blackwell MLPerf training recordsBenchmark result
The companies submitted the largest NVIDIA Blackwell MLPerf Training v5.0 cluster and reported record Llama 3.1 405B training performance.
May 20252
May 16, 2025
Unauthorized Grok prompt modification disclosedSafety report
xAI disclosed that an unauthorized system-prompt change caused Grok to inject unrelated claims about South African racial politics; it published prompts and announced stronger review and monitoring controls.
May 14, 2025
AlphaEvolve introducedResearch paper
Google DeepMind introduces AlphaEvolve, a Gemini-powered evolutionary coding agent for algorithm discovery and optimization.
April 20251
Apr 2, 2025
CoreWeave posts first cloud NVIDIA GB200 MLPerf inference resultsBenchmark result
CoreWeave submitted the first cloud-provider MLPerf Inference v5.0 results for NVIDIA GB200, reporting more than 800 tokens per second on Llama 3.1 405B.
March 20252
Mar 27, 2025
Tracing the thoughts of a large language model publishedResearch paper
Anthropic published interpretability research tracing internal mechanisms underlying multilingual reasoning, planning, and unfaithful explanations.
Mar 2025
Kimina-Prover project introducedResearch paper
Moonshot introduced the Kimina-Prover line of open models for formal theorem proving.
February 20257
Feb 19, 2025
Muse world and action model introducedResearch paper
Microsoft Research and Xbox introduce Muse, a generative world and action model trained on gameplay sequences.
Feb 11, 2025
MASK belief-alignment benchmark releasedBenchmark result
Scale published MASK as part of its safety research program for testing whether models knowingly state beliefs inconsistent with their internal representations.
Feb 11, 2025
FORTRESS benchmark releasedBenchmark result
Scale released FORTRESS to evaluate frontier-model safeguards against dual-use national-security and public-safety risks.
Feb 10, 2025
Anthropic Economic Index launchedResearch paper
Anthropic launched the Economic Index to measure how AI is being used across occupations and tasks.
Feb 7, 2025
PARTNR benchmark and dataset releasedResearch paper
Meta released PARTNR, a benchmark, dataset and planning model for human-robot collaboration on household tasks.
Feb 6, 2025
Brain2Qwerty research publishedResearch paper
Meta researchers introduced a non-invasive deep-learning system for decoding typed sentences from EEG and MEG brain recordings.
Feb 3, 2025
Constitutional Classifiers research publishedResearch paper
Anthropic published and tested Constitutional Classifiers, a safeguard approach designed to resist broad classes of jailbreaks.
January 20255
Jan 23, 2025
Humanity's Last Exam results and benchmark releasedBenchmark result
Scale and the Center for AI Safety published Humanity’s Last Exam, assembled from nearly 1,000 contributors across more than 500 institutions in 50 countries.
Jan 20, 2025
DeepSeek R1 reasoning family launchesResearch paper
DeepSeek introduced its reinforcement-learning reasoning model family and technical report.
Jan 16, 2025
MatterGen research publishedResearch paper
Microsoft Research publishes MatterGen, a generative model that designs stable inorganic materials conditioned on target properties.
Jan 2025
Kimi k1 reasoning model disclosedResearch paper
Moonshot described k1 as a long-context reinforcement-learning reasoning system.
2025
TRIBE research model recognized at Algonauts 2025Benchmark result
Meta's first TRIBE brain-response model formed the foundation for its award-winning Algonauts 2025 entry; Meta's later TRIBE v2 announcement identifies the original model as 2025 work.
December 20244
Dec 26, 2024
DeepSeek V3 family launchesResearch paper
DeepSeek introduced its 671B-parameter V3 mixture-of-experts architecture.
Dec 18, 2024
Alignment faking research publishedResearch paper
Anthropic and Redwood Research published evidence of alignment-faking behavior in a controlled Claude 3 Opus experiment.
Dec 4, 2024
BioEmu generative protein model introducedResearch paper
Microsoft Research introduces BioEmu for efficiently sampling diverse protein conformational ensembles.
Dec 4, 2024
Genie 2 world model introducedResearch paper
Google DeepMind introduces Genie 2, capable of generating diverse playable 3D environments from a prompt image.
November 20241
Nov 2024
Kimi k0-math research model disclosedResearch paper
Moonshot disclosed k0-math as an experimental reinforcement-learning model for mathematical reasoning.
October 20243
Oct 18, 2024
Janus multimodal family introducedResearch paper
DeepSeek introduced the Janus family for unified multimodal understanding and generation.
Oct 18, 2024
Spirit LM introducedResearch paper
Meta released research on a multimodal language model that freely mixes text and speech.
Oct 4, 2024
Movie Gen model family introducedResearch paper
Meta introduced foundation models for video generation, personalized video, editing and synchronized audio.
September 20241
Sep 2024
EnigmaEval benchmark releasedBenchmark result
Scale and research collaborators introduced EnigmaEval, a multimodal reasoning benchmark built from novel puzzle-competition problems.
August 20241
Aug 2024
MultiChallenge benchmark releasedBenchmark result
Scale released MultiChallenge to measure model performance on realistic multi-turn conversation problems involving memory, instruction retention, editing, and self-consistency.
July 20241
Jul 2, 2024
Meta 3D Gen introducedResearch paper
Meta introduced a research pipeline for generating high-quality 3D assets with text and texture control.
June 20242
Jun 13, 2024
Microsoft delays broad Recall rollout for security reviewSafety report
Microsoft changes Recall from broad Copilot+ PC availability to a Windows Insider preview while strengthening security and privacy protections.
Jun 1, 2024
DeepSeek-Prover research line introducedResearch paper
DeepSeek introduced its formal theorem-proving model research line.
May 20244
May 23, 2024
Golden Gate Claude interpretability experiment releasedResearch paper
Anthropic demonstrated feature amplification in Claude by making a discovered Golden Gate Bridge feature unusually active.
May 21, 2024
Mapping the mind of a large language model publishedResearch paper
Anthropic published a large-scale mechanistic interpretability study mapping millions of concepts inside Claude 3 Sonnet.
May 20, 2024
Aurora foundation model introducedResearch paper
Microsoft Research introduces Aurora, a foundation model for weather and atmospheric forecasting across multiple environmental domains.
May 16, 2024
Chameleon research model introducedResearch paper
Meta introduced an early-fusion multimodal foundation model capable of generating and reasoning over interleaved text and images.
April 20241
Apr 2, 2024
Many-shot jailbreaking research publishedResearch paper
Anthropic published research showing how long-context models can be jailbroken using many demonstrations and proposed mitigations.
March 20244
Mar 13, 2024
SIMA generalist game agent introducedResearch paper
Google DeepMind introduces SIMA, an instructable agent that follows natural-language directions across multiple 3D environments.
Mar 11, 2024
DeepSeek-VL family introducedResearch paper
DeepSeek introduced an open vision-language model family for real-world multimodal understanding.
Mar 2024
Mooncake serving architecture releasedResearch paper
Moonshot released Mooncake, a KV-cache-centric architecture for disaggregated large-model serving.
Mar 2024
Moonshot publishes open research infrastructureResearch paper
Moonshot began publishing open systems research for efficient long-context model training and serving.
February 20243
Feb 26, 2024
Genie generative interactive environments introducedResearch paper
Google DeepMind introduces Genie, a generative interactive environment model trained from unlabelled internet videos.
Feb 13, 2024
Aya open-science initiative produces first modelResearch paper
Cohere Labs' Aya initiative culminated in its first open multilingual model and dataset work spanning 101 languages.
Feb 5, 2024
DeepSeekMath family introducedResearch paper
DeepSeek introduced its open mathematical reasoning research line.
January 20241
Jan 12, 2024
AMIE diagnostic-dialogue research introducedResearch paper
Google Research introduces AMIE, a research AI system optimized for diagnostic medical dialogue.
December 20232
Dec 11, 2023
Audiobox research model introducedResearch paper
Meta introduced Audiobox, a foundation research model for generating and editing speech, voices and sound effects.
Dec 7, 2023
Purple Llama initiative launchedSafety report
Meta launched Purple Llama, an open trust-and-safety initiative for generative AI.
November 20235
Nov 29, 2023
GNoME predicts new stable materialsResearch paper
Google DeepMind reports that GNoME identified 2.2 million new crystal structures, including hundreds of thousands of stable candidates.
Nov 20, 2023
Orca 2 research models releasedResearch paper
Microsoft Research releases Orca 2 in 7B and 13B sizes, training smaller models to use different reasoning strategies.
Nov 16, 2023
Emu Edit research releasedResearch paper
Meta published Emu Edit for instruction-based image editing with precise task control.
Nov 16, 2023
Emu Video research releasedResearch paper
Meta published Emu Video, a factorized text-to-video generation method.
Nov 14, 2023
GraphCast research publishedResearch paper
Google DeepMind publishes GraphCast, a machine-learning model for fast, accurate medium-range global weather forecasts.
October 20231
Oct 19, 2023
IBM Research unveils NorthPole AI chipResearch paper
IBM Research published a memory-centric neural inference chip designed for high energy efficiency.
September 20231
Sep 19, 2023
AlphaMissense classifies missense variantsResearch paper
Google DeepMind publishes AlphaMissense predictions for the effects of millions of human missense variants.
June 20234
Jun 20, 2023
Phi-1 research model introducedResearch paper
Microsoft Research introduces Phi-1, showing that curated textbook-quality data can produce strong code and reasoning performance in a compact model.
Jun 20, 2023
Kosmos-2 grounded multimodal model introducedResearch paper
Microsoft Research introduces Kosmos-2, linking generated language to object regions in images through grounding.
Jun 16, 2023
Voicebox introducedResearch paper
Meta introduced Voicebox, a generative speech model capable of editing, sampling and cross-lingual speech generation.
Jun 5, 2023
Orca research model introducedResearch paper
Microsoft Research introduces Orca, a 13-billion-parameter model trained from rich explanation traces produced by larger models.
May 20231
May 9, 2023
Claude's Constitution publishedResearch paper
Anthropic published the principles used to train and guide Claude's behavior through Constitutional AI.
February 20231
Feb 27, 2023
Kosmos-1 research publishedResearch paper
Microsoft Research introduces Kosmos-1, a multimodal large language model capable of perceiving visual inputs and following language instructions.
January 20231
Jan 26, 2023
MusicLM research introducedResearch paper
Google Research introduces MusicLM, a hierarchical text-to-music generation model for high-fidelity audio.
December 20222
Dec 26, 2022
Med-PaLM research introducedResearch paper
Google Research presents Med-PaLM, adapting PaLM to medical question answering and clinical knowledge benchmarks.
Dec 15, 2022
Constitutional AI research publishedResearch paper
Anthropic published Constitutional AI, a method for training a harmless assistant through AI feedback guided by explicit principles.
October 20222
Oct 18, 2022
IBM Research unveils Artificial Intelligence UnitResearch paper
IBM Research disclosed an AI accelerator built for efficient enterprise inference.
Oct 7, 2022
AudioLM research introducedResearch paper
Google Research introduces AudioLM for generating long-term coherent, high-fidelity audio from discrete representations.
September 20221
Sep 29, 2022
Make-A-Video introducedResearch paper
Meta presented Make-A-Video, a system for generating short videos from text prompts.
August 20221
Aug 2, 2022
AlexaTM 20B research publishedResearch paper
Amazon researchers published a 20-billion-parameter multilingual sequence-to-sequence teacher model.
July 20221
Jul 14, 2022
Make-A-Scene introducedResearch paper
Meta introduced Make-A-Scene, a multimodal generative method that gives creators scene-level control over images.
May 20221
May 23, 2022
Imagen research introducedResearch paper
Google Research introduces Imagen, a text-to-image diffusion model emphasizing photorealism and language understanding.
April 20221
Apr 4, 2022
PaLM 540B introducedResearch paper
Google Research introduces the 540-billion-parameter Pathways Language Model and reports strong few-shot performance across hundreds of tasks.
October 20211
Oct 14, 2021
Ego4D dataset launchedResearch paper
Facebook AI and university partners introduced Ego4D, a large egocentric video dataset and benchmark suite.
July 20211
Jul 6, 2021
ERNIE 3.0 research releasedResearch paper
Baidu published its unified framework for knowledge-enhanced language understanding and generation.
May 20211
May 21, 2021
wav2vec-U publishedResearch paper
Facebook AI introduced unsupervised speech recognition using unpaired speech audio and text.
December 20201
Dec 23, 2020
MuZero research published in NatureResearch paper
DeepMind publishes MuZero, which plans using a learned model without being told the environment's underlying rules.
November 20201
Nov 30, 2020
AlphaFold 2 solves the CASP14 protein-folding challengeBenchmark result
AlphaFold 2 demonstrates accuracy competitive with experimental structures at CASP14, marking a major protein-folding breakthrough.
June 20201
Jun 20, 2020
wav2vec 2.0 publishedResearch paper
Facebook AI introduced wav2vec 2.0 for self-supervised speech learning with limited labeled data.

