Imagining in 360° decouples humanoid visual search into an Imaginator predicting spatial priors and an Actor using them to improve search efficiency without costly trajectory annotations.
TAC generates temporally grounded audio captions via synthetic training, reducing hallucinations and outperforming competitors in detection and dense captioning; cascading it with LLMs achieves state-of-the-art audio and audio-visual reasoning.
Inline Critic uses learnable tokens to critique frozen image-editing models at intermediate layers, steering hidden states during the forward pass to achieve state-of-the-art results.
StreamGaze introduces a benchmark for evaluating gaze-guided temporal and proactive reasoning in streaming videos, revealing large performance gaps between state-of-the-art MLLMs and humans.
Conditioning diffusion models on multimodal large language models with VAE identity conditioning and dual-layer aggregation improves subject-driven generation by balancing semantics with identity preservation.
RoPEMover manipulates diffusion transformer position embeddings to move objects with depth-aware 3D geometry, preserving identity, occlusions, and shadows with minimal real data.
Auteur introduces a human-centric DSL and virtual director to generate language-driven camera trajectories for human-centered video synthesis, outperforming existing methods on framing metrics.