UniT unifies human and humanoid actions via visual anchoring into shared latent tokens for scalable policy learning and world modeling. It achieves state-of-the-art data efficiency, zero-shot transfer, and cross-embodiment dynamics alignment.
AnyMo introduces OmniHuMo dataset with 5,000 hours of multimodal motion data and proposes a masked modeling framework for scalable any-modality conditional motion synthesis with flexible spatial and stylistic control.
PhotoFlow uses a Director-Reviewer-Reflector agent for closed-loop camera search to generate language-conditioned virtual photographs in arbitrary 3D scenes, outperforming baselines on quality, alignment, and success rate.
BALTO applies balanced token-level credit assignment to mitigate LLM hallucinations by redistributing probability from unsupported to faithful content, outperforming response-level methods on faithfulness benchmarks.
SimCT recovers lost cross-tokenizer supervision by comparing multi-token continuations in on-policy distillation, improving reasoning and code generation over exact token matching.
Advanced optimizers weaken multi-task learning because instant gradients barely affect updates, so the APT framework with adaptive momentum and Muon direction preservation improves results across four datasets.
TokenHD trains token-level hallucination detectors via scalable synthetic data and importance weighting, with small models outperforming larger reasoning models and scaling consistently.
FaithfulFaces improves identity-preserving video generation via pose-shared identity alignment and achieves state-of-the-art consistency across pose changes and occlusions.
Balanced Fine-Tuning uses dual-scale token and sequence reweighting targeting dense epistemic uncertainty to align LLMs with biomedical knowledge, improving reasoning and sparse-reward RL over standard fine-tuning.
A game-theoretic framework predicts and steers LLM populations via Nash equilibrium analysis, deriving closed-form alignments that prevent political exclusion and guide socially desirable outcomes.
NAVA proposes native audio-visual alignment with an Align-then-Fuse MMDiT architecture for joint audio-video generation, achieving superior synchronization, video quality, and timbre control with 6.3B parameters.
OpenSearch-VL introduces an open-source recipe training multimodal deep search agents via curated data, diverse tools, and multi-turn fatal-aware GRPO, achieving over 10-point benchmark gains comparable to proprietary models.
FDEP integrates frozen visual foundation model representations into infrared small target detection via semantic alignment fusion and implicit self-distillation, achieving state-of-the-art accuracy with no inference overhead.