Gen-Searcher trains a search-augmented image generation agent via supervised and reinforcement learning, yielding about 16-point gains on knowledge-intensive benchmarks.
RWML learns action-conditioned world models for LLM agents via self-supervised sim-to-real alignment, outperforming direct task-success RL by up to 6.9 points without expert data.
PARE combines structure-aware width pruning and timestep-conditioned adaptive depth routing to cut video diffusion compute while preserving generation quality.
CELLO predicts single-cell spatial transcriptomics from histology via grid sampling and distance-decay cross-attention, achieving 14x faster inference than DeepSpot2Cell without upstream segmentation.
Confidence-based decoding achieves ε-accurate diffusion language model sampling in Õ(H(X₀)/ε) iterations by adaptively unmasking tokens until cumulative entropy exceeds a threshold.
MultiTalk introduces 57.6k hours of synthetic multi-party bilingual dialogue data and MultiTalkBench for long-form full-duplex evaluation, training a model that sustains coherent extended multi-party English-Chinese conversation and outperforms open-source baselines.
DriveDreamer-Policy unifies depth generation, video prediction, and motion planning via geometry-aware world representations, achieving 89.2 PDMS on Navsim v1 and 88.7 EPDMS on v2.
EO-WM is a diffusion transformer that forecasts satellite imagery via physically structured weather conditioning and improves vegetation-decline prediction accuracy by up to 7.8% on new diagnostic benchmarks.
MemDLM augments diffusion language model training via bi-level optimization with parametric memory, improving convergence, long-context representations, and needle retrieval.
OpenWebRL enables open online RL for visual web agents, with a 4B model reaching 67% Online-Mind2Web and 64% DeepShop success using minimal initialization data.
Direct product flow matching decouples radial and angular dynamics via constant-speed geodesic transport and hidden-state conditioning for state-of-the-art few-shot vision-language adaptation.
BitDance is an autoregressive image generator that predicts binary visual tokens via a diffusion head and next-patch decoding, achieving state-of-the-art FID with far fewer parameters and much faster inference.
xHC expands Transformer hyper-connections beyond four streams via sparse updates and temporal augmentation, improving scaling efficiency. It boosts 18B MoE downstream scores by 4.0 points over mHC with lower compute and reduced memory traffic via xHC-Flash.
SMTL replaces sequential reasoning with parallel evidence acquisition for efficient long-horizon agentic search, achieving state-of-the-art results on multiple benchmarks with far fewer reasoning steps.
WorldMemArena evaluates multimodal agent memory through an action-world loop, showing writing and storage improvements do not guarantee performance and harness-based memory remains costly and unreliable.
ReFPO adds explicit reflow regularization to flow matching policy gradients, stabilizing training and enabling high-fidelity one-step inference that matches multi-step performance across control tasks.
A visual-native harness with an image bank and on-policy data evolution improves multimodal deep search agents, raising Qwen3-VL-8B to 39.0% average and surpassing Gemini-2.5 Pro.
OpenCoF introduces a 17K video reasoning dataset and Wan-CoF model that improves chain-of-frame reasoning via diverse temporal supervision and reasoning tokens.
UniClawBench introduces a capability-driven benchmark evaluating proactive agents via 400 real-world tasks with live Docker evaluation and multi-turn feedback.
AnchorWorld improves egocentric world simulation via full-body interaction supervision and anchor-view customization with consistent spatio-temporal dynamics.
OpenSearch-VL introduces an open-source recipe training multimodal deep search agents via curated data, diverse tools, and multi-turn fatal-aware GRPO, achieving over 10-point benchmark gains comparable to proprietary models.
FDEP integrates frozen visual foundation model representations into infrared small target detection via semantic alignment fusion and implicit self-distillation, achieving state-of-the-art accuracy with no inference overhead.
FocusDepth uses spatially-aligned multi-scale prompt fusion to boost target-region depth accuracy and sharp boundaries while preserving global geometry, outperforming global baselines on FDE-Bench.
ONE-SHOT factorizes compositional video generation via spatial-decoupled motion injection and hybrid context integration, achieving fine-grained human-environment control and minute-level consistency without 3D alignment.
TextPro-SLM minimizes speech-text modality gap by feeding prosody-aware text inputs to LLMs, cutting the gap at 3B and 7B scales with only ~1,000 hours of audio.
OQRC partitions calibration losses into ordered bins to tightly upper-bound quantile risk with finite-sample guarantees converging at rate O_p(n^{-1/2}).
Astra enhances vision-language model spatial reasoning by letting agents generate imagined simulator views via RL, improving MMSI-Bench scores over direct answering.
DoAtlas-1 introduces causal compilation to convert medical evidence into executable causal estimands, achieving 98.5% canonicalization accuracy and 80.5% query executability across 1,445 effect kernels.
Seg3DParts uses segmentation-grounded generation with structured cross-part interaction to produce controllable, coherent part-level 3D meshes from single images while introducing the PartObjectNet dataset.
DirectUV generates UV textures via diffusion with surface-aware positional encoding that attains 3D coherence across seams and improves occluded regions.
StraTA introduces trajectory-level strategies into agentic reinforcement learning via hierarchical rollout training, improving long-horizon decision-making and reaching 93.1% on ALFWorld.
Self-evolving LLMs can misevolve into dangerous entities, and anchoring a small safety circuit during evolution preserves safety with minimal capability loss.
Fixed-path attribution uniquely requires Aumann-Shapley line integrals, while transport-geodesic paths via minimized kinetic action yield more stable, structured explanations.
Quantum estimators for heavy-tailed noise enable QNSGD and QPSGD to find ε-stationary or optimal solutions with poly(√d, ε) oracle queries, beating classical lower bounds in low dimensions.
PLANING decouples geometry and appearance via explicit primitives and neural Gaussians for fast, high-quality streaming monocular 3D reconstruction with reduced redundancy.
EcoGym benchmarks long-horizon LLM economic planning across open-source environments, revealing no single model dominates and exposing strategic and execution suboptimalities.
IRR-Drive uses adaptive multimodal text and BEV reflection to self-correct driving intentions before trajectory generation, achieving state-of-the-art NAVSIM results.
FASTER accelerates real-time flow vision-language-action models via horizon-aware sampling that compresses immediate-action denoising into one step, slashing reaction latency on dynamic robot tasks.
LaST-R1 uses reinforcement learning with adaptive latent reasoning to optimize robotic action policies, achieving 99.9% success on LIBERO and up to 22.5% real-world gains.
ViCO minimizes vision tokens via consistency training across compression ratios, cutting tokens up to 50% while preserving capabilities through semantic-based routing.
CAM uses continuous knowledge graph extraction and adaptive multi-method querying to answer entity-centric video questions, improving accuracy by up to 23 points over baselines.
CausalMix frames data mixture optimization as causal inference to dynamically estimate optimal mixtures via conditional average treatment effects, improving LLM performance without retraining proxy models.