Gen-Searcher trains a search-augmented image generation agent via supervised and reinforcement learning, yielding about 16-point gains on knowledge-intensive benchmarks.
RWML learns action-conditioned world models for LLM agents via self-supervised sim-to-real alignment, outperforming direct task-success RL by up to 6.9 points without expert data.
PARE combines structure-aware width pruning and timestep-conditioned adaptive depth routing to cut video diffusion compute while preserving generation quality.
CELLO predicts single-cell spatial transcriptomics from histology via grid sampling and distance-decay cross-attention, achieving 14x faster inference than DeepSpot2Cell without upstream segmentation.
Confidence-based decoding achieves ε-accurate diffusion language model sampling in Õ(H(X₀)/ε) iterations by adaptively unmasking tokens until cumulative entropy exceeds a threshold.
MultiTalk introduces 57.6k hours of synthetic multi-party bilingual dialogue data and MultiTalkBench for long-form full-duplex evaluation, training a model that sustains coherent extended multi-party English-Chinese conversation and outperforms open-source baselines.
DriveDreamer-Policy unifies depth generation, video prediction, and motion planning via geometry-aware world representations, achieving 89.2 PDMS on Navsim v1 and 88.7 EPDMS on v2.