SIGMA uses semantic feature differencing with instruction-guided spatial priors to generate manipulation masks from edited images, producing a 1.1M training set that improves detectors by +18.34% F1.
HumanoidArena benchmarks egocentric hierarchical whole-body learning via seven leg-critical tasks, finding policies solve diverse interactions but cross-tracker transfer remains fragile.