DiscoverPhysics benchmarks LLM agents on simulated worlds with non-standard physics, finding frontier models pass only half and fail at uncovering latent structure.
Semantic uncertainty measures answer disagreement rather than reliability, as valid answers vary and repeated errors appear certain; a bias-uncertainty decomposition separates variability from systematic error to improve evaluation.