RWML learns action-conditioned world models for LLM agents via self-supervised sim-to-real alignment, outperforming direct task-success RL by up to 6.9 points without expert data.
OATS dynamically generates training-stage-specific synthetic data guided by valuable samples via diffusion, consistently outperforming static augmentation for time series foundation models.
P2T uses reference patches as privileged supervision to curate shorter, grounded agent trajectories via bi-objective optimization, improving SWE-bench Pass@1 by up to 10.8 points with ~15% lower inference cost.
OpenWebRL enables open online RL for visual web agents, with a 4B model reaching 67% Online-Mind2Web and 64% DeepShop success using minimal initialization data.
Standard reconstruction in latent action models encodes exogenous future state, but endogenous representations and action supervision prevent this failure.
Unsupervised reward modeling via web document prefix-suffix preference learning improves RewardBench accuracy up to 7.7 points and matches supervised baselines without human annotations.
MorphoHELM benchmarks microscopy representation methods across batch effects, finding classic computer vision strategies outperform deep learning across settings and revealing trade-offs between models.
RAW-Dream disentangles world models from task data by using task-agnostic pre-trained dynamics and VLM rewards to fine-tune VLAs entirely in zero-shot imagination with verified rollouts.
CodeScaler uses a reward model to scale code LLM training and inference without test cases, improving benchmarks by up to 14.64 points and cutting latency tenfold.
This paper models parallel inference-time reasoning via particle filtering, deriving non-asymptotic guarantees and fundamental limits for sequential Monte Carlo with process reward models.
Information-flow control secures AI agents via Fides, a planner with deterministic confidentiality and integrity tracking that completes diverse AgentDojo tasks with guarantees.
A framework generates population-aligned personas from social media via quality filtering, importance sampling, and task-specific adaptation, reducing bias in LLM social simulations.
Video diffusion transformers encode physical plausibility linearly in latent states around 81% accuracy despite lacking predictive training, emerging inside the denoiser rather than the VAE.
GAP fixes visual-latent reasoning instability via feature, context, and capacity alignment, improving Qwen2.5-VL 7B perception and reasoning performance.
Monroe is a molecular foundation model pre-trained on 81 million molecules that uses in-context TabPFN prediction to achieve state-of-the-art bioassay activity prediction, especially on activity cliffs.
Vermeer is an autoregressive generative model that predicts protein localization microscopy from sequences and cell landmarks, enabling zero-shot transfer across imaging conditions.
NextLat adds latent self-prediction to transformers, theoretically converging to belief states and empirically improving world modeling, reasoning, and inference speed.