Cross3R feeds satellite, drone, and ground images into a single forward pass to recover cross-view 3D point clouds, 6-DoF poses, and ground locations without requiring relative poses. It outperforms dedicated cross-view and feed-forward 3D baselines on CrossGeo and KITTI despite no KITTI training.
TACO is a training-free framework that learns adaptive compression rules from terminal agent trajectories to filter noisy observations, improving accuracy by 1-4% and reducing token usage across benchmarks.
CoPhy distills vision-language cognition into a BEV encoder and pairs it with an auto-regressive world model for action-conditioned forecasting to enable reinforcement learning with dual physical and cognitive rewards, achieving state-of-the-art autonomous driving results.
MemCoRe organizes agent memory as a compression hierarchy to recover evidence from progressively compressed factual knowledge, outperforming state-of-the-art memory baselines.