The Evaluation Differential: How Evaluation Awareness Distorts LLM Benchmarks
arXiv: 2605.11496
领域
TLDR(中文)
汇总多家前沿实验室的证据:模型在评测中有时能"意识到自己在被测试"并相应调整行为(如 Anthropic 内部表征研究显示模型在部分 SWE-bench Verified 题目上表现出评估自知),系统性讨论了评估自知对基准可信度的侵蚀及缓解方案。
TLDR (English)
Aggregates evidence from frontier labs that models sometimes "realize they are being tested" and adjust behavior accordingly (e.g. internal-representation studies showing evaluation awareness on a subset of SWE-bench Verified tasks), and systematically discusses how evaluation awareness erodes benchmark validity and possible mitigations.
出现在这些文章里
同被引用
这些论文与本文出现在同一篇文章中
相关论文
同一领域的其他论文