IAMFlow is a training-free identity-aware memory framework that tracks persistent entities across prompts to generate consistent long narrative videos, achieving best benchmark scores and faster inference.
SRL-MPC integrates reinforcement-learned parameter updates with shape-aware model predictive control via geometric separation features to navigate dense heterogeneous robot crowds safely and adaptively.
Timeflies jointly infers whether future observations exist and predicts their values, outperforming methods that assume future observation times are known.
MORA rewrites prompts to expand multi-dimensional reward diversity and breaks the safety-helpfulness trade-off, improving sequential single-preference alignment by up to 12.4% and simultaneous alignment by 4.6%.
SkeMex improves medical agents via self-evolving skill memory that distills reusable procedural knowledge, governs retention by utility, and outperforms memory-based agents across clinical tasks.
EO-WM is a diffusion transformer that forecasts satellite imagery via physically structured weather conditioning and improves vegetation-decline prediction accuracy by up to 7.8% on new diagnostic benchmarks.
SUGAR converts human videos into humanoid loco-manipulation skills via automated priors, physics refinement, and policy distillation, scaling with video data and enabling zero-shot real-world transfer.
Standard video backbone readouts suppress patch-level temporal dynamics needed to detect AI-generated videos; a lightweight velocity-gated patch profiling readout reaches 95.28 AUC on frozen backbones.
ITO improves image-text pretraining via multi-view cross-modal alignment and discarded training-time fusion, beating CLIP at 100M-1B scale on classification and retrieval.
StableVQ decouples encoder-decoder and codebook training via Dynamic STE, Region VQ Loss, and independent schedules to stabilize VQ tokenizers and boost utilization and reconstruction.
TMPO replaces scalar reward maximization with trajectory-level reward distribution matching via Softmax Trajectory Balance, improving diffusion alignment diversity by 9.1% while avoiding reward hacking and mode collapse.
PointForward reconstructs driving scenes via world-space 3D queries and scene graphs, achieving state-of-the-art feedforward results with explicit cross-view and instance consistency.
Next Forcing uses multi-chunk prediction to accelerate convergence 2.3x, boost high-frame-rate accuracy 93.1%, and double inference speed for world models.
DoAtlas-1 introduces causal compilation to convert medical evidence into executable causal estimands, achieving 98.5% canonicalization accuracy and 80.5% query executability across 1,445 effect kernels.
LaMo extracts self-supervised latent motion priors from unlabeled videos via motion drift loss and prior guidance, improving physical consistency in video diffusion without external supervision.
ReMind trains video diffusion transformers to use cache memory for evolving hidden states across interruptions via memory-oriented curricula and PM-RoPE, achieving best STEVO-Bench scores without catastrophic forgetting.
Falcon-X maps heterogeneous time series variates into a unified latent prototype space using diff-attention and latent entity attention to enable cross-variate modeling and zero-shot structural transfer with strong forecasting results.
MDPBench introduces a 3,400-image multilingual document parsing benchmark across 17 languages revealing open-source models suffer severe performance drops on photographed and non-Latin script documents.
K12-KGraph introduces a curriculum-aligned K-12 knowledge graph, benchmark, and training data showing current LLMs achieve under 57 percent accuracy on curriculum cognition and that graph-guided supervision outperforms generic instruction tuning.
LatentUM unifies modalities in a shared latent space to enable efficient interleaved cross-modal reasoning and generation, achieving state-of-the-art visual planning and self-reflective generation results.
FASTER accelerates real-time flow vision-language-action models via horizon-aware sampling that compresses immediate-action denoising into one step, slashing reaction latency on dynamic robot tasks.
A structured spectral propagator in a latent space improves long-horizon PDE forecasting stability, reducing relative L2 errors by up to 48.9% versus state-of-the-art methods.