跳转到内容

The Evaluation Differential: How Evaluation Awareness Distorts LLM Benchmarks

作者: The Evaluation Differential Team (2026)

arXiv: 2605.11496

领域

评估安全

TLDR(中文)

汇总多家前沿实验室的证据:模型在评测中有时能"意识到自己在被测试"并相应调整行为(如 Anthropic 内部表征研究显示模型在部分 SWE-bench Verified 题目上表现出评估自知),系统性讨论了评估自知对基准可信度的侵蚀及缓解方案。

TLDR (English)

Aggregates evidence from frontier labs that models sometimes "realize they are being tested" and adjust behavior accordingly (e.g. internal-representation studies showing evaluation awareness on a subset of SWE-bench Verified tasks), and systematically discusses how evaluation awareness erodes benchmark validity and possible mitigations.

出现在这些文章里

同被引用

这些论文与本文出现在同一篇文章中

相关论文

同一领域的其他论文