Skip to content

LiveCodeBench

Canonical page on the main site: sinoaihub.com/benchmarks/livecodebench

Type

benchmark

Key facts

  • Task type: Competitive-programming code generation plus self-repair, code execution and test-output prediction
  • Dataset size: Continuously growing; initially 300+ problems from LeetCode, AtCoder and Codeforces contests (from May 2023 onward, updated over time)
  • Evaluation method: Models solve fresh contest problems released after a cutoff date; graded by executing code against hidden test cases
  • Scoring: Pass@1 accuracy (% of problems solved correctly on the first submission)
  • Recorded evaluations: 0

Description

A contamination-controlled coding benchmark that continuously collects fresh problems from LeetCode, AtCoder and Codeforces contests, plus self-repair, code-execution and test-output-prediction tasks.

Evaluations

No evaluations recorded — no tracked Chinese model publishes a LiveCodeBench score as of 2026-09-29.

Methodology

Task type: Competitive-programming code generation plus self-repair, code execution and test-output prediction

Dataset size: Continuously growing; initially 300+ problems from LeetCode, AtCoder and Codeforces contests (from May 2023 onward, updated over time)

Evaluation method: Models solve fresh contest problems released after a cutoff date; graded by executing code against hidden test cases

Scoring: Pass@1 accuracy (% of problems solved correctly on the first submission)

Contamination notes: Problems are drawn from contests after a fixed cutoff date specifically to avoid contamination; the benchmark keeps adding new problems to stay ahead of training leakage.

Relevant Models

No tracked model currently publishes a LiveCodeBench score. Natural candidates in the database include DeepSeek-V4.1-Flash, Kimi K2.7 Code and GLM-5.3-Flash.

Limitations

Pass@1 depends on sampling temperature and prompting, and on the execution sandbox. Newer problems are added continuously, so a score is always tied to a specific snapshot. No Chinese model in the China AI Hub database currently publishes a LiveCodeBench score, so this page records no evaluations.

Verification Status

verified

Benchmark Changes

No documented benchmark changes on record as of 2026-09-29.

Last Verified

2026-09-29

Source history

No documented source-change events located as of 2026-09-29.

Sources

evidence_id source_name source_url source_type published verified confidence conflict
src-benchmarks-livecodebench-1 Jain et al. — LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code https://arxiv.org/abs/2403.07974 Literature 2024-03 2026-09-29 high —
src-benchmarks-livecodebench-2 LiveCodeBench website https://livecodebench.github.io/ Official documentation — 2026-09-29 high —
src-benchmarks-livecodebench-3 LiveCodeBench repository https://github.com/LiveCodeBench/LiveCodeBench Official documentation — 2026-09-29 high —