Event Cascade Pruning uses event-camera motion cues to prune video tokens in first-person spatial reasoning, improving accuracy by 1.31 points with 80% fewer tokens and 1.89x speedup.
A unified framework decomposes LLM alignment dynamics into competing rebound and driving forces, explaining reversal and faster re-alignment via rehearsal priming.
UniGraphLM proposes a unified graph language model that multi-domain multi-task aligns GNN representations to LLMs via adaptive alignment for cross-domain generalization.
MM-IssueLoc benchmarks multimodal repository-level issue localization using visual evidence across 652 instances, showing current systems achieve under 39% file accuracy and text-only scores do not transfer.
Deriving an alignment imprint from LLM preference tuning, LAPD detects AI-generated text with 45.82% relative gains over baselines via statistically guaranteed preference discrepancy.
DriveDreamer-Policy unifies depth generation, video prediction, and motion planning via geometry-aware world representations, achieving 89.2 PDMS on Navsim v1 and 88.7 EPDMS on v2.
MyoChallenge 2025 benchmarks musculoskeletal sports control via simulated table tennis and soccer tasks, advancing agile motor algorithms across 70 teams.
VETime unifies temporal and visual modalities via fine-grained alignment and dynamic fusion for zero-shot time-series anomaly detection, outperforming state-of-the-art models with lower overhead.
ECHO-2 is a distributed RL framework that overlaps rollout generation, dissemination, and training with bounded policy staleness to improve cost efficiency while preserving rewards.
A visual-native harness with an image bank and on-policy data evolution improves multimodal deep search agents, raising Qwen3-VL-8B to 39.0% average and surpassing Gemini-2.5 Pro.