Selective layer looping improves masked diffusion model training efficiency and reasoning performance via depth scaling without added parameters and flexible inference compute scaling.
EHR-ReasonCon introduces a reasoning-intensive benchmark for clinical note-table consistency verification, and EHR-Inspector achieves state-of-the-art results via LLM-based verification with table exploration.
Condition-dependent source distributions for flow matching improve text-to-image generation via variance regularization and directional alignment, accelerating convergence up to 3x in FID.
SELFCI uses complementary self-distillation to decouple information suppression from task resolution, improving contextual integrity without degrading utility.
AMUSE integrates Muon's rapid bulk progress with Schedule-Free averaging via time-varying interpolation to suppress oscillations, requiring no learning rate schedules and improving training efficiency across vision and LLM tasks.
TRQAM introduces trust-region Q-adjoint matching with adaptive path-space KL control via projected dual descent, enabling stable off-policy flow-policy fine-tuning and achieving 68% success on OGBench.
Kernel Discovery uses an LLM-driven evolutionary framework to search broad kernel spaces for high-dimensional Bayesian optimization, achieving average rank 1.2 out of 17.
STRATA predicts lipid nanoparticle transfection by aligning molecular structure and composition ratio representations to model component interactions. It improves accuracy and generalizes to unseen molecules and ratios.
MotionGrounder enables multi-object motion transfer via a diffusion transformer with flow-based motion signals, object-caption alignment loss, and a new object grounding score. It outperforms baselines in multi-object controllable video generation.
Multi-view relational distillation improves vision-language model spatial reasoning by distilling cross-view patch similarities rather than features, preserving language alignment with minimal overhead.
MASF redesigns score-based filtering with a measurement-aware forward process that transforms states toward measurements, yielding accurate likelihood scores, stronger assimilation under sparse nonlinear observations, and up to 28.2x faster inference.
ViTeX-Bench introduces a 387-video benchmark and evaluation protocol for high-fidelity video scene text editing, finding that accuracy, temporal stability, and edit locality remain hard to balance together.
CDM amortizes twisted SMC for discrete diffusion by learning a twist function via contrastive samples, adding under 5% overhead while outperforming baselines on text, DNA, protein, and LLM tasks.
Attention sinks in Omni-LLMs serve as global representation biases rather than redundant heads, and the proposed OutRo decoding method improves video reasoning with minimal overhead.
Bellman residual minimization for policy optimization achieves stable convergence with function approximation but lacks extensive study; foundational control results are established.
CRePE encodes tokens as depth-aware distributions along curved unified-camera rays to unify camera control, lens geometry, and external geometry guidance in video generation.
TriProRep pretrains structure-aware protein representations via joint amino-acid, backbone, and full-atom token recovery, improving homodimer co-folding, interaction prediction, and monomer structure prediction over sequence-only models.
SpatialClaw uses a stateful Python kernel with step-wise code execution to enable flexible spatial reasoning, achieving 59.9% average accuracy across 20 benchmarks.
VideoLLMs lose temporal representations across layers, so Temporal Activation Injection reinforces fading divergence at inference to improve video reasoning without training.
ES-Merging estimates merging coefficients from embedding space signals to combine biological multimodal models, improving cross-modal reasoning and single-modal knowledge preservation.
POP enables context-conditioned online structural pruning of foundation models via coarse-to-fine partitioned masking without offline calibration or retraining, improving accuracy with lower latency.
GeoCore-9B is a 9-billion-parameter diffusion model trained from scratch on global earth observation data that conditions generation on geospatial metadata and achieves state-of-the-art visual fidelity and geographic accuracy.
VocalCoachBench benchmarks audio-language models on expert singing feedback, revealing they identify broad vocal issues but fall below baselines on fine-grained diagnosis and strict alignment.
PISCO enables precise video instance insertion via sparse keyframe control while preserving dynamics, achieving monotonic gains with added signals and outperforming editing baselines.
A unified sign-language model with privacy-preserving keypoint inputs and sliding perceiver aggregation achieves state-of-the-art translation and alignment on BSL and generalizes to ASL.