SuperCLUE
Canonical page on the main site: sinoaihub.com/benchmarks/superclue
Type
benchmark
Key facts
- Task type: Chinese multi-part evaluation spanning language understanding/generation, professional knowledge, agent capability and safety
- Dataset size: Multi-part rolling evaluation across four ability quadrants and 12 base capabilities; no fixed single question count (periodic leaderboard releases)
- Evaluation method: Combines objective multiple-choice questions with multi-turn open-ended questions and agent/tool-use tasks; scored by the CLUE team rather than self-reported
- Scoring: Composite total score (总分) with sub-scores per quadrant (e.g. OPEN multi-turn, open questions, objective questions)
- Recorded evaluations: 0
Description
A Chinese comprehensive evaluation system for large models, organised around four ability quadrants — language understanding and generation, professional knowledge, agent capability and safety — refined into 12 base capabilities.
Evaluations
No evaluations recorded — no tracked Chinese model publishes a SuperCLUE score as of 2026-09-29.
Methodology
Task type: Chinese multi-part evaluation spanning language understanding/generation, professional knowledge, agent capability and safety
Dataset size: Multi-part rolling evaluation across four ability quadrants and 12 base capabilities; no fixed single question count (periodic leaderboard releases)
Evaluation method: Combines objective multiple-choice questions with multi-turn open-ended questions and agent/tool-use tasks; scored by the CLUE team rather than self-reported
Scoring: Composite total score (总分) with sub-scores per quadrant (e.g. OPEN multi-turn, open questions, objective questions)
Contamination notes: SuperCLUE maintains public and hidden test sets and runs its own evaluation, reducing reliance on vendor self-reported scores.
Relevant Models
No tracked model currently publishes a SuperCLUE score. Natural candidates in the database include Doubao Seed 2.1 Pro, GLM-5.3 and Qwen3.8-Max.
Limitations
The composite score aggregates heterogeneous sub-tasks, so a single total hides capability-specific strengths. Leaderboard snapshots are time-stamped and not comparable across dates. No Chinese model in the China AI Hub database currently publishes a SuperCLUE score, so this page records no evaluations.
Verification Status
verified
Benchmark Changes
No documented benchmark changes on record as of 2026-09-29.
Last Verified
2026-09-29
Source history
No documented source-change events located as of 2026-09-29.
Sources
| evidence_id | source_name | source_url | source_type | published | verified | confidence | conflict |
|---|---|---|---|---|---|---|---|
| src-benchmarks-superclue-1 | SuperCLUE: A Comprehensive Chinese Large Language Model Benchmark | https://arxiv.org/abs/2307.15020 | Literature | 2023-07 | 2026-09-29 | high | — |
| src-benchmarks-superclue-2 | SuperCLUE official repository (CLUEbenchmark) | https://github.com/CLUEbenchmark/SuperCLUE | Official documentation | — | 2026-09-29 | high | — |
| src-benchmarks-superclue-3 | SuperCLUE official website | https://www.superclueai.com/ | Official documentation | — | 2026-09-29 | high | — |