Benchmark evaluating whether language models can maintain instructions, user information, editing state, and self-consistency across multi-turn conversations.
- Aug 1, 2024
- —
- —
- —
- —
- 2
Releases
Mar 23, 2026MultiChallenge evaluation pipeline updatedFeature update
Scale Labs updated MultiChallenge with a stronger judge model, refined tasks, and refreshed results to improve agreement with expert ratings.
Aug 2024MultiChallenge benchmark releasedBenchmark result
Scale released MultiChallenge to measure model performance on realistic multi-turn conversation problems involving memory, instruction retention, editing, and self-consistency.

