Releases
123 matching — model launches, API releases, product updates, financial events and news.
June 20251
Jun 4, 2025
CoreWeave, NVIDIA, and IBM set Blackwell MLPerf training recordsBenchmark result
The companies submitted the largest NVIDIA Blackwell MLPerf Training v5.0 cluster and reported record Llama 3.1 405B training performance.
May 20252
May 16, 2025
Unauthorized Grok prompt modification disclosedSafety report
xAI disclosed that an unauthorized system-prompt change caused Grok to inject unrelated claims about South African racial politics; it published prompts and announced stronger review and monitoring controls.
May 14, 2025
AlphaEvolve introducedResearch paper
Google DeepMind introduces AlphaEvolve, a Gemini-powered evolutionary coding agent for algorithm discovery and optimization.
April 20251
Apr 2, 2025
CoreWeave posts first cloud NVIDIA GB200 MLPerf inference resultsBenchmark result
CoreWeave submitted the first cloud-provider MLPerf Inference v5.0 results for NVIDIA GB200, reporting more than 800 tokens per second on Llama 3.1 405B.
March 20252
Mar 27, 2025
Tracing the thoughts of a large language model publishedResearch paper
Anthropic published interpretability research tracing internal mechanisms underlying multilingual reasoning, planning, and unfaithful explanations.
Mar 2025
Kimina-Prover project introducedResearch paper
Moonshot introduced the Kimina-Prover line of open models for formal theorem proving.
February 20257
Feb 19, 2025
Muse world and action model introducedResearch paper
Microsoft Research and Xbox introduce Muse, a generative world and action model trained on gameplay sequences.
Feb 11, 2025
MASK belief-alignment benchmark releasedBenchmark result
Scale published MASK as part of its safety research program for testing whether models knowingly state beliefs inconsistent with their internal representations.
Feb 11, 2025
FORTRESS benchmark releasedBenchmark result
Scale released FORTRESS to evaluate frontier-model safeguards against dual-use national-security and public-safety risks.
Feb 10, 2025
Anthropic Economic Index launchedResearch paper
Anthropic launched the Economic Index to measure how AI is being used across occupations and tasks.
Feb 7, 2025
PARTNR benchmark and dataset releasedResearch paper
Meta released PARTNR, a benchmark, dataset and planning model for human-robot collaboration on household tasks.
Feb 6, 2025
Brain2Qwerty research publishedResearch paper
Meta researchers introduced a non-invasive deep-learning system for decoding typed sentences from EEG and MEG brain recordings.
Feb 3, 2025
Constitutional Classifiers research publishedResearch paper
Anthropic published and tested Constitutional Classifiers, a safeguard approach designed to resist broad classes of jailbreaks.
January 20255
Jan 23, 2025
Humanity's Last Exam results and benchmark releasedBenchmark result
Scale and the Center for AI Safety published Humanity’s Last Exam, assembled from nearly 1,000 contributors across more than 500 institutions in 50 countries.
Jan 20, 2025
DeepSeek R1 reasoning family launchesResearch paper
DeepSeek introduced its reinforcement-learning reasoning model family and technical report.
Jan 16, 2025
MatterGen research publishedResearch paper
Microsoft Research publishes MatterGen, a generative model that designs stable inorganic materials conditioned on target properties.
Jan 2025
Kimi k1 reasoning model disclosedResearch paper
Moonshot described k1 as a long-context reinforcement-learning reasoning system.
2025
TRIBE research model recognized at Algonauts 2025Benchmark result
Meta's first TRIBE brain-response model formed the foundation for its award-winning Algonauts 2025 entry; Meta's later TRIBE v2 announcement identifies the original model as 2025 work.
December 20244
Dec 26, 2024
DeepSeek V3 family launchesResearch paper
DeepSeek introduced its 671B-parameter V3 mixture-of-experts architecture.
Dec 18, 2024
Alignment faking research publishedResearch paper
Anthropic and Redwood Research published evidence of alignment-faking behavior in a controlled Claude 3 Opus experiment.
Dec 4, 2024
BioEmu generative protein model introducedResearch paper
Microsoft Research introduces BioEmu for efficiently sampling diverse protein conformational ensembles.
Dec 4, 2024
Genie 2 world model introducedResearch paper
Google DeepMind introduces Genie 2, capable of generating diverse playable 3D environments from a prompt image.
November 20241
Nov 2024
Kimi k0-math research model disclosedResearch paper
Moonshot disclosed k0-math as an experimental reinforcement-learning model for mathematical reasoning.
October 20242
Oct 18, 2024
Janus multimodal family introducedResearch paper
DeepSeek introduced the Janus family for unified multimodal understanding and generation.
Oct 18, 2024
Spirit LM introducedResearch paper
Meta released research on a multimodal language model that freely mixes text and speech.

