IAarxiv.org·hace 2 días·actualizado automáticamente

EvalDetectBench: A Benchmark for Measuring Evaluation Awareness in Frontier Language Models

arXiv:2609.01611v1 Announce Type: new Abstract: Frontier large language models can often recognize when they are being evaluated, a capability known as evaluation awareness. If models behave differently in evaluations than in deployment, this undermines the validity of evaluation results, which are a crucial component of current AI safety frameworks. We introduce EvalDetectBench, an open pipeline and benchmark for measuring evaluation awareness that works with any Inspect-compatible evaluation,

Debate en la comunidad

Todavía no hay comentarios. Sé el primero en opinar.