C2R generates realistic, temporally consistent urban crowd videos from coarse 3D simulations via a neural renderer guided by text and a synthetic-real domain-hedging strategy.
MoMHa treats LLM harness design as multi-objective search over accuracy, safety, and token cost, outperforming baselines across 17 domains via joint-reward optimization.
MotionGrounder enables multi-object motion transfer via a diffusion transformer with flow-based motion signals, object-caption alignment loss, and a new object grounding score. It outperforms baselines in multi-object controllable video generation.
PixelDense aligns pixel diffusion with frozen dense-prediction teachers via separate semantic and geometric projection streams and orthogonality penalties, improving GenEval to 0.8093 and training speed by 1.23x.
Using summed last-k encoder layers and combining RAE with REPA, RAEv2 achieves state-of-the-art gFID of 1.06 in 80 epochs with 10x faster convergence and free guidance.
TAC generates temporally grounded audio captions via synthetic training, reducing hallucinations and outperforming competitors in detection and dense captioning; cascading it with LLMs achieves state-of-the-art audio and audio-visual reasoning.
Inline Critic uses learnable tokens to critique frozen image-editing models at intermediate layers, steering hidden states during the forward pass to achieve state-of-the-art results.
A dataset of multi-aspect human visual similarity judgments benchmarks vision-language models and yields the TPIPS metric, which aligns with human perception and enables text-guided image retrieval and generative evaluation.
StreamGaze introduces a benchmark for evaluating gaze-guided temporal and proactive reasoning in streaming videos, revealing large performance gaps between state-of-the-art MLLMs and humans.
Activation patching's natural indirect effect embeds hidden interaction effects between components, which cause conditional importance to be invisible or inflated, explain faithfulness instability, scale with activation distance, and diagnose when greedy component ranking misses combinatorial mechan
Conditioning diffusion models on multimodal large language models with VAE identity conditioning and dual-layer aggregation improves subject-driven generation by balancing semantics with identity preservation.
ParetoSlider trains one diffusion model with continuous preference weights to approximate the full Pareto front, enabling inference-time navigation of trade-offs between conflicting generative goals without retraining.
UNITE unifies tokenization and latent diffusion via a shared generative encoder, enabling single-stage joint training from scratch without adversarial losses or pretrained encoders to reach near state-of-the-art FID scores.