SP-Mem decouples memory utility from private-value exposure via isolated storage and consent-based retrieval, improving personalization while reducing unnecessary privacy exposure.
Cross3R feeds satellite, drone, and ground images into a single forward pass to recover cross-view 3D point clouds, 6-DoF poses, and ground locations without requiring relative poses. It outperforms dedicated cross-view and feed-forward 3D baselines on CrossGeo and KITTI despite no KITTI training.
ATI-VLA aligns predictive observations and actions in a shared discrete codebook, then adaptively injects predictive latents into action decoding, achieving state-of-the-art robotic manipulation with faster convergence.
ReSCUE enables simultaneous sign language translation on continuous unsegmented video via inference-aware training, stabilized re-translation, and online sentence commitment, achieving low-latency quality near offline oracles.
SyncWorld learns action-visual mappings via visual calibration episodes to serve as zero-shot simulators across unseen robot settings without retraining.
StableHand estimates world-space dual-hand motion from egocentric video via quality-aware flow matching, cutting W-MPJPE by 20-25% over baselines on occluded benchmarks.
A neural-behavioral framework decodes natural whole-body monkey movement from large-scale epidural cortical signals via an autoregressive model without physical constraints.