FocusDepth uses spatially-aligned multi-scale prompt fusion to boost target-region depth accuracy and sharp boundaries while preserving global geometry, outperforming global baselines on FDE-Bench.
METIS is a vision-language-action model pretrained on multi-source egocentric data that achieves the highest success rate across six real-world dexterous manipulation tasks.
EA-WM projects robotic actions into camera-aligned visual fields and fuses them via event-aware attention to preserve spatial geometry and interaction dynamics, achieving state-of-the-art results on WorldArena.
Evo-Depth is a lightweight 0.9-billion-parameter vision-language-action model using implicit depth encoding from RGB to improve spatial manipulation without extra sensors. It achieves top benchmark performance with minimal GPU memory and highest inference speed among compared methods.