C-Eval
Canonical page on the main site: sinoaihub.com/benchmarks/c-eval
Type
benchmark
Key facts
- Task type: Multiple-choice knowledge and reasoning questions across 52 disciplines (humanities, science, engineering)
- Dataset size: 13,948 multiple-choice questions across 52 disciplines; four difficulty levels (middle school, high school, college, professional); plus the C-Eval Hard subset
- Evaluation method: Closed-book multiple-choice; models answer exam-style questions spanning 52 disciplines at four difficulty levels
- Scoring: Accuracy (% correct); reported as an overall average and per-discipline / per-level breakdowns
- Recorded evaluations: 0
Description
A Chinese evaluation suite of 13,948 multiple-choice questions across 52 disciplines and four difficulty levels, for measuring advanced knowledge and reasoning of foundation models in a Chinese context.
Evaluations
No evaluations recorded — no tracked Chinese model publishes a C-Eval score as of 2026-09-29.
Methodology
Task type: Multiple-choice knowledge and reasoning questions across 52 disciplines (humanities, science, engineering)
Dataset size: 13,948 multiple-choice questions across 52 disciplines; four difficulty levels (middle school, high school, college, professional); plus the C-Eval Hard subset
Evaluation method: Closed-book multiple-choice; models answer exam-style questions spanning 52 disciplines at four difficulty levels
Scoring: Accuracy (% correct); reported as an overall average and per-discipline / per-level breakdowns
Contamination notes: The complete C-Eval test set was released to the community in July 2025, so models trained after that date may have seen the answers.
Relevant Models
No tracked model currently publishes a C-Eval score. Natural candidates in the database include Qwen3.8-Max, GLM-5.3 and Kimi K3.
Limitations
Multiple-choice format cannot assess free-form generation, open-ended reasoning or factual grounding. Scores on C-Eval Hard and the full set are not interchangeable. No Chinese model in the China AI Hub database currently publishes a C-Eval score, so this page records no evaluations.
Verification Status
verified
Benchmark Changes
No documented benchmark changes on record as of 2026-09-29.
Last Verified
2026-09-29
Source history
No documented source-change events located as of 2026-09-29.
Sources
| evidence_id | source_name | source_url | source_type | published | verified | confidence | conflict |
|---|---|---|---|---|---|---|---|
| src-benchmarks-c-eval-1 | Huang et al. — C-Eval: A Multi-Level Multi-Discipline Chinese Evaluation Suite for Foundation Models | https://arxiv.org/abs/2305.08322 | Literature | 2023-05 | 2026-09-29 | high | — |
| src-benchmarks-c-eval-2 | C-Eval official repository (HKUST-NLP) | https://github.com/hkust-nlp/ceval | Official documentation | — | 2026-09-29 | high | — |
| src-benchmarks-c-eval-3 | C-Eval benchmark website | https://cevalbenchmark.com/ | Official documentation | — | 2026-09-29 | high | — |