OpenWebRL enables open online RL for visual web agents, with a 4B model reaching 67% Online-Mind2Web and 64% DeepShop success using minimal initialization data.
GPFlow learns generalized Poisson process rate functions for variable-length protein generation, improving designability and recovering length distributions without fixed-length constraints.
Visual Sparse Steering trains sparse autoencoders on frozen CLIP activations to build label-free steering vectors that improve zero-shot classification by up to 4.12 percent via centroid-deviation steering with reconstruction-error gating.
A clustering-based divergence method measures gaps between real and simulated user behaviors, finding large, family-dependent discrepancies reducible by combining complementary simulators.
Lean4Agent uses Lean4 to formally model and verify agent workflows, with verified workflows outperforming failing ones by 11.94% and LeanEvolve improving SWE performance by 7.47%.
CoT-Guard, a 4B-parameter chain-of-thought monitor, detects hidden code-generation objectives via SFT and RL, outperforming larger models including GPT-5.
SkillOS uses RL to train a skill curator that updates an external SkillRepo from experience, improving self-evolving agents across reasoning and multi-turn tasks.
Exact posterior score estimation derives closed-form posterior scores for linear Gaussian inverse problems, enabling efficient training and sampling that outperforms baselines with far fewer evaluations.
Trajectory self-distillation trains few-step diffusion language models to match full-step trajectories, mitigating factorization error to enable fast, high-quality parallel decoding.
ReToken introduces one learnable retrieval token that selects sparse visual tokens from long contexts, improving vision-language models by up to 13.4 points on visual retrieval while fitting on a single GPU.
This paper proposes Laws of Reasoning (LoRe), a framework formalizing reasoning compute and accuracy laws, plus LoRe-Bench showing models lack compositionality; enforcing compute-law compositionality via finetuning improves reasoning performance.
LangFlow closes the continuous-discrete language-modeling gap via flow matching and a learnable noise schedule, matching discrete diffusion perplexity and exceeding autoregressive zero-shot results on four benchmarks.
OrchRM uses self-supervised orchestration-level reward modeling to train multi-agent orchestrators, cutting token usage by 10x and boosting accuracy up to 8%.
MemReward propagates rewards through a heterogeneous rollout graph to enable LLM reinforcement learning using only 20% ground-truth labels and achieves over 96% of oracle performance.
Multiplicative tâtonnement uses local per-chore excess-demand updates to converge to competitive equilibrium in CCH disutility chores markets. It achieves O(1/ε²) approximate equilibrium rates with better constants and runs substantially faster in practice.