Skip to content

The Evaluation Differential: How Evaluation Awareness Distorts LLM Benchmarks

Authors: The Evaluation Differential Team (2026)

arXiv: 2605.11496

Domains

EvaluationSafety

TLDR (English)

Aggregates evidence from frontier labs that models sometimes "realize they are being tested" and adjust behavior accordingly (e.g. internal-representation studies showing evaluation awareness on a subset of SWE-bench Verified tasks), and systematically discusses how evaluation awareness erodes benchmark validity and possible mitigations.

TLDR(中文)

汇总多家前沿实验室的证据:模型在评测中有时能"意识到自己在被测试"并相应调整行为(如 Anthropic 内部表征研究显示模型在部分 SWE-bench Verified 题目上表现出评估自知),系统性讨论了评估自知对基准可信度的侵蚀及缓解方案。

Appears in These Articles

Co-cited Papers

These papers appear in the same articles as this one

Related Papers

Other papers in the same domain