UniT unifies human and humanoid actions via visual anchoring into shared latent tokens for scalable policy learning and world modeling. It achieves state-of-the-art data efficiency, zero-shot transfer, and cross-embodiment dynamics alignment.
AnyMo introduces OmniHuMo dataset with 5,000 hours of multimodal motion data and proposes a masked modeling framework for scalable any-modality conditional motion synthesis with flexible spatial and stylistic control.
PhotoFlow uses a Director-Reviewer-Reflector agent for closed-loop camera search to generate language-conditioned virtual photographs in arbitrary 3D scenes, outperforming baselines on quality, alignment, and success rate.
BALTO applies balanced token-level credit assignment to mitigate LLM hallucinations by redistributing probability from unsupported to faithful content, outperforming response-level methods on faithfulness benchmarks.
SimCT recovers lost cross-tokenizer supervision by comparing multi-token continuations in on-policy distillation, improving reasoning and code generation over exact token matching.
Advanced optimizers weaken multi-task learning because instant gradients barely affect updates, so the APT framework with adaptive momentum and Muon direction preservation improves results across four datasets.
TokenHD trains token-level hallucination detectors via scalable synthetic data and importance weighting, with small models outperforming larger reasoning models and scaling consistently.
FaithfulFaces improves identity-preserving video generation via pose-shared identity alignment and achieves state-of-the-art consistency across pose changes and occlusions.
Balanced Fine-Tuning uses dual-scale token and sequence reweighting targeting dense epistemic uncertainty to align LLMs with biomedical knowledge, improving reasoning and sparse-reward RL over standard fine-tuning.
A game-theoretic framework predicts and steers LLM populations via Nash equilibrium analysis, deriving closed-form alignments that prevent political exclusion and guide socially desirable outcomes.
NAVA proposes native audio-visual alignment with an Align-then-Fuse MMDiT architecture for joint audio-video generation, achieving superior synchronization, video quality, and timbre control with 6.3B parameters.
OpenSearch-VL introduces an open-source recipe training multimodal deep search agents via curated data, diverse tools, and multi-turn fatal-aware GRPO, achieving over 10-point benchmark gains comparable to proprietary models.
FDEP integrates frozen visual foundation model representations into infrared small target detection via semantic alignment fusion and implicit self-distillation, achieving state-of-the-art accuracy with no inference overhead.
Delta-Adapter extracts a semantic delta from single image pairs to train exemplar-based editors without paired examples, improving accuracy and generalization.
PRECISE introduces an SDE-consistent stochastic sampler balancing exploration and stability for RL post-training of flow-matching models, enabling faster, more stable reward optimization with significantly reduced training time.
Direct corpus interaction uses terminal tools to search raw corpora directly, bypassing fixed retrieval interfaces and substantially outperforming sparse, dense, and reranking baselines on agentic search benchmarks.
Psy-CoT decomposes role-playing reasoning into psychology-grounded steps, and RAPO uses profile-token mutual information to weight gradients, improving fidelity and out-of-distribution generalization over supervised fine-tuning.
ReDiff reframes vision-language diffusion as active refining via error-correction training and online self-correction loops, breaking error cascades to improve coherence, factual accuracy, and parallel generation.
Hydra-X unifies image and video tokenization in one vision transformer via causal temporal attention and hierarchical compression, achieving strong unified understanding and generation performance.
Flow-DPPO replaces PPO ratio clipping with exact KL divergence constraints for flow matching models, improving reward, stability, and multi-objective alignment.