Gen-Searcher trains a search-augmented image generation agent via supervised and reinforcement learning, yielding about 16-point gains on knowledge-intensive benchmarks.
RWML learns action-conditioned world models for LLM agents via self-supervised sim-to-real alignment, outperforming direct task-success RL by up to 6.9 points without expert data.
PARE combines structure-aware width pruning and timestep-conditioned adaptive depth routing to cut video diffusion compute while preserving generation quality.
CELLO predicts single-cell spatial transcriptomics from histology via grid sampling and distance-decay cross-attention, achieving 14x faster inference than DeepSpot2Cell without upstream segmentation.
Confidence-based decoding achieves ε-accurate diffusion language model sampling in Õ(H(X₀)/ε) iterations by adaptively unmasking tokens until cumulative entropy exceeds a threshold.
MultiTalk introduces 57.6k hours of synthetic multi-party bilingual dialogue data and MultiTalkBench for long-form full-duplex evaluation, training a model that sustains coherent extended multi-party English-Chinese conversation and outperforms open-source baselines.
DriveDreamer-Policy unifies depth generation, video prediction, and motion planning via geometry-aware world representations, achieving 89.2 PDMS on Navsim v1 and 88.7 EPDMS on v2.
EO-WM is a diffusion transformer that forecasts satellite imagery via physically structured weather conditioning and improves vegetation-decline prediction accuracy by up to 7.8% on new diagnostic benchmarks.
MemDLM augments diffusion language model training via bi-level optimization with parametric memory, improving convergence, long-context representations, and needle retrieval.
OpenWebRL enables open online RL for visual web agents, with a 4B model reaching 67% Online-Mind2Web and 64% DeepShop success using minimal initialization data.
Direct product flow matching decouples radial and angular dynamics via constant-speed geodesic transport and hidden-state conditioning for state-of-the-art few-shot vision-language adaptation.
BitDance is an autoregressive image generator that predicts binary visual tokens via a diffusion head and next-patch decoding, achieving state-of-the-art FID with far fewer parameters and much faster inference.
xHC expands Transformer hyper-connections beyond four streams via sparse updates and temporal augmentation, improving scaling efficiency. It boosts 18B MoE downstream scores by 4.0 points over mHC with lower compute and reduced memory traffic via xHC-Flash.
SMTL replaces sequential reasoning with parallel evidence acquisition for efficient long-horizon agentic search, achieving state-of-the-art results on multiple benchmarks with far fewer reasoning steps.
WorldMemArena evaluates multimodal agent memory through an action-world loop, showing writing and storage improvements do not guarantee performance and harness-based memory remains costly and unreliable.
ReFPO adds explicit reflow regularization to flow matching policy gradients, stabilizing training and enabling high-fidelity one-step inference that matches multi-step performance across control tasks.
A visual-native harness with an image bank and on-policy data evolution improves multimodal deep search agents, raising Qwen3-VL-8B to 39.0% average and surpassing Gemini-2.5 Pro.
OpenCoF introduces a 17K video reasoning dataset and Wan-CoF model that improves chain-of-frame reasoning via diverse temporal supervision and reasoning tokens.
UniClawBench introduces a capability-driven benchmark evaluating proactive agents via 400 real-world tasks with live Docker evaluation and multi-turn feedback.
AnchorWorld improves egocentric world simulation via full-body interaction supervision and anchor-view customization with consistent spatio-temporal dynamics.
OpenSearch-VL introduces an open-source recipe training multimodal deep search agents via curated data, diverse tools, and multi-turn fatal-aware GRPO, achieving over 10-point benchmark gains comparable to proprietary models.
FDEP integrates frozen visual foundation model representations into infrared small target detection via semantic alignment fusion and implicit self-distillation, achieving state-of-the-art accuracy with no inference overhead.