Skip to content

GAIA

Canonical page on the main site: sinoaihub.com/benchmarks/gaia

Type

benchmark

Key facts

  • Task type: Real-world question answering requiring reasoning, web browsing, multi-modality and tool use
  • Dataset size: 466 questions across three difficulty levels; answers to 300 of them are held out for the leaderboard
  • Evaluation method: Questions are conceptually simple for humans but require tool use; graded by exact-match against a single ground-truth answer
  • Scoring: Exact-match accuracy (% correct); a question is passed only if the final answer matches exactly
  • Recorded evaluations: 0

Description

A benchmark for general AI assistants: 466 real-world questions requiring reasoning, multi-modality handling, web browsing and tool use, where humans score 92% and early frontier assistants far lower.

Evaluations

No evaluations recorded — no tracked Chinese model publishes a GAIA score as of 2026-09-29.

Methodology

Task type: Real-world question answering requiring reasoning, web browsing, multi-modality and tool use

Dataset size: 466 questions across three difficulty levels; answers to 300 of them are held out for the leaderboard

Evaluation method: Questions are conceptually simple for humans but require tool use; graded by exact-match against a single ground-truth answer

Scoring: Exact-match accuracy (% correct); a question is passed only if the final answer matches exactly

Contamination notes: 300 of the 466 questions are retained privately for leaderboard evaluation, reducing training-set leakage.

Relevant Models

No tracked model currently publishes a GAIA score. Natural candidates in the database include DeepSeek-V4-Pro and Qwen3.8-Max.

Limitations

Exact-match scoring is strict and can understate near-correct answers. Results depend on the tool/search scaffolding a model is given, so scores are not comparable across different setups. No Chinese model in the China AI Hub database currently publishes a GAIA score, so this page records no evaluations.

Verification Status

verified

Benchmark Changes

No documented benchmark changes on record as of 2026-09-29.

Last Verified

2026-09-29

Source history

No documented source-change events located as of 2026-09-29.

Sources

evidence_id source_name source_url source_type published verified confidence conflict
src-benchmarks-gaia-1 Mialon et al. — GAIA: a benchmark for General AI Assistants https://arxiv.org/abs/2311.12983 Literature 2023-11 2026-09-29 high —
src-benchmarks-gaia-2 GAIA benchmark (Hugging Face) https://huggingface.co/gaia-benchmark Official documentation — 2026-09-29 high —