Skip to content

IFEval

Canonical page on the main site: sinoaihub.com/benchmarks/ifeval

Type

benchmark

Key facts

  • Task type: Instruction following on verifiable natural-language constraints (length, format, keyword counts, etc.)
  • Dataset size: 25 types of verifiable instructions; approximately 500 prompts, each containing one or more verifiable instructions
  • Evaluation method: Prompts carry rule-checkable instructions; a program verifies whether each constraint was satisfied
  • Scoring: Strict accuracy and prompt-level / instruction-level accuracy (fraction of instructions followed)
  • Recorded evaluations: 0

Description

Instruction-Following Eval: a benchmark of verifiable instructions ("write over 400 words", "mention a keyword at least 3 times") that checks whether a model actually obeys precise natural-language constraints, scored by rule-checking rather than an LLM judge.

Evaluations

No evaluations recorded — no tracked Chinese model publishes an IFEval score as of 2026-09-29.

Methodology

Task type: Instruction following on verifiable natural-language constraints (length, format, keyword counts, etc.)

Dataset size: 25 types of verifiable instructions; approximately 500 prompts, each containing one or more verifiable instructions

Evaluation method: Prompts carry rule-checkable instructions; a program verifies whether each constraint was satisfied

Scoring: Strict accuracy and prompt-level / instruction-level accuracy (fraction of instructions followed)

Contamination notes: Rule-checked verification (not an LLM judge) reduces evaluator bias; the verifiable-instruction format is not tied to a single knowledge cutoff.

Relevant Models

No tracked model currently publishes an IFEval score. Natural candidates in the database include Qwen3.8-Max, GLM-5.3 and DeepSeek-V4-Pro.

Limitations

IFEval measures only mechanical, rule-checkable instruction following — it does not capture semantic quality, helpfulness or correctness of the response content. No Chinese model in the China AI Hub database currently publishes an IFEval score, so this page records no evaluations.

Verification Status

verified

Benchmark Changes

No documented benchmark changes on record as of 2026-09-29.

Last Verified

2026-09-29

Source history

No documented source-change events located as of 2026-09-29.

Sources

evidence_id source_name source_url source_type published verified confidence conflict
src-benchmarks-ifeval-1 Zhou et al. — Instruction-Following Evaluation for Large Language Models (IFEval) https://arxiv.org/abs/2311.07911 Literature 2023-11 2026-09-29 high —
src-benchmarks-ifeval-2 IFEval — Google Research (instruction_following_eval) https://github.com/google-research/google-research/tree/master/instruction_following_eval Official documentation — 2026-09-29 high —