CMMLU
Canonical page on the main site: sinoaihub.com/benchmarks/cmmlu
Type
benchmark
Key facts
- Task type: Chinese multiple-choice knowledge and reasoning questions across 67 subjects
- Dataset size: 67 subjects from elementary to advanced professional level; multiple-choice with 4 options and a single correct answer; a 5-question development set and a 100+ question test set per subject
- Evaluation method: Multiple-choice; evaluated in zero-shot and five-shot settings, with and without chain-of-thought
- Scoring: Accuracy (% correct); the random baseline is 25%
- Recorded evaluations: 0
Description
A comprehensive Chinese benchmark measuring massive multitask language understanding across 67 subjects, from STEM and humanities to China-specific knowledge such as driving rules.
Evaluations
No evaluations recorded — no tracked Chinese model publishes a CMMLU score as of 2026-09-29.
Methodology
Task type: Chinese multiple-choice knowledge and reasoning questions across 67 subjects
Dataset size: 67 subjects from elementary to advanced professional level; multiple-choice with 4 options and a single correct answer; a 5-question development set and a 100+ question test set per subject
Evaluation method: Multiple-choice; evaluated in zero-shot and five-shot settings, with and without chain-of-thought
Scoring: Accuracy (% correct); the random baseline is 25%
Contamination notes: CMMLU maintainers verify API-only models for data contamination before listing them on the leaderboard; questions include China-specific answers less common in English training data.
Relevant Models
No tracked model currently publishes a CMMLU score. Natural candidates in the database include Qwen3.8-Max, DeepSeek-V4-Pro and GLM-5.3.
Limitations
Multiple-choice format measures recognition and reasoning within fixed options, not free-form generation or grounding. Zero-shot and five-shot scores are not directly comparable. No Chinese model in the China AI Hub database currently publishes a CMMLU score, so this page records no evaluations.
Verification Status
verified
Benchmark Changes
No documented benchmark changes on record as of 2026-09-29.
Last Verified
2026-09-29
Source history
No documented source-change events located as of 2026-09-29.
Sources
| evidence_id | source_name | source_url | source_type | published | verified | confidence | conflict |
|---|---|---|---|---|---|---|---|
| src-benchmarks-cmmlu-1 | Li et al. — CMMLU: Measuring massive multitask language understanding in Chinese | https://arxiv.org/abs/2306.09212 | Literature | 2023-06 | 2026-09-29 | high | — |
| src-benchmarks-cmmlu-2 | CMMLU official repository | https://github.com/haonan-li/CMMLU | Official documentation | — | 2026-09-29 | high | — |
| src-benchmarks-cmmlu-3 | CMMLU dataset card (Hugging Face) | https://huggingface.co/datasets/haonan-li/cmmlu | Official documentation | — | 2026-09-29 | high | — |