SLVMBench evaluates video-LLMs on learning skills from long video streams and applying them in real time, revealing substantial performance degradation.
HOMIE unifies inter- and intra-subject video personalization via multimodal guidance and reference embeddings, achieving state-of-the-art human-object interaction fidelity.
HACRL enables heterogeneous agents to share verified rollouts during collaborative on-policy training and execute independently at inference, with HACPO improving all agents by 3.6% over baselines at half the rollout cost.
ActWorld extends interactive world models to object interaction via a 100K dataset and hierarchical action-aware memory, improving fidelity over navigation-only baselines.
NLAC trains LLM agents with a natural-language generative critic for off-policy learning, yielding richer feedback and more stable, data-efficient training than policy gradients in long-horizon tasks.