4DCodeBench: Benchmarking Agents on Inverse Graphics of Dynamic Scenes
4DCodeBench benchmarks agents reconstructing dynamic scenes from video as executable graphics code, finding strong static models fail at complex dynamics.
Published Oct 2, 2026▲ 28 on Hugging FaceCode ★ 89arXiv ↗

Only vote on papers you've read. Sign in with GitHub to vote.
4DCodeBench offers a rigorous, needed testbed for dynamic inverse graphics via code, but its missing baselines, dataset scale, and precise failure modes leave its frontier-model gap more asserted than fully validated.
Abstract
We introduce 4DCodeBench, a benchmark for 4D inverse graphics through code generation, in which agents reconstruct dynamic scenes from video as executable graphics programs. To accomplish this, agents must translate visual observations into compact representations of scene structure and dynamics, by implementing abstractions such as physical simulations to reproduce complex behavior. To evaluate this capability, we curate a set of real-world videos and construct synthetic scenes spanning diverse physical phenomena, including deformation, fluid flow, and fracture. We perform extensive benchmarking of frontier models, finding that strong static reconstruction capabilities do not yet translate into reliable reconstruction of complex dynamics. 4DCodeBench provides a testbed for tracking progress toward agents that can interpret the dynamics of the world through code. Our benchmark is available at https://github.com/4DCodeBench/4DCodeBench