Skip to content
AIMarketCap

MultiChallenge

Product
Scale AIPart of SEALactive

Benchmark evaluating whether language models can maintain instructions, user information, editing state, and self-consistency across multi-turn conversations.

Launched
Aug 1, 2024
Models
—
Modality
—
Open source
—
API available
—
Releases
2

Releases

2
Mar 23, 2026MultiChallenge evaluation pipeline updatedMultiChallenge

Scale Labs updated MultiChallenge with a stronger judge model, refined tasks, and refreshed results to improve agreement with expert ratings.

Aug 2024MultiChallenge benchmark releasedMultiChallenge

Scale released MultiChallenge to measure model performance on realistic multi-turn conversation problems involving memory, instruction retention, editing, and self-consistency.