Models with evaluation meta-knowledge about benchmark structures score safer via implicit behavioral shifts, confounding safety assessments independently of explicit awareness.
ZendoWorld evaluates AI agents on active visual rule induction and finds high prediction accuracy does not imply rule recovery, with VLM agents proposing near-uninformative experiments.