PAGER closes the semantic-execution gap for point-precise geometric GUI control via dependency-structured planning and pixel-level execution, achieving 4.1x higher task success than general baselines.
RACES recursively composes verifiable environments as LEGO bricks to scale RL reasoning training, boosting model performance on unseen benchmarks with far fewer base environments.
IRIS unifies self-play fine-tuning via adjustable Rényi divergence with adaptive schedules, surpassing supervised fine-tuning with fewer annotations across benchmarks.
PAPO-VLA improves vision-language-action reliability by identifying planning actions via action variation and trajectory outcomes, weighting them by causal importance in policy optimization, and boosting benchmark performance.
PSD accelerates diffusion LLM inference via adaptive parallel unmasking and multi-depth speculative drafts with hierarchical verification, achieving up to 5.5x tokens per pass with near-greedy accuracy.
ViDiC introduces a video difference captioning task and ViDiC-1K benchmark that reveals large multimodal models struggle with fine-grained comparative video perception.
GraDE uses a graph diffusion estimator to score subgraph typicality and discovers large-scale neural architecture motifs with up to 30x higher median frequency than sampling methods.
RLSD combines RLVR and self-distillation, using token-level policy differences for update magnitudes and environmental feedback for directions, improving convergence and stability.
Hyper-spherical quantization decouples semantics from magnitude via angular routing to prevent codebook collapse, enabling scalable discrete representation autoencoders with full codebook usage and high-fidelity reconstruction.
A shared gradient-interaction scoring rule unifies parameter and data selection for LLM fine-tuning via DualSFT, improving joint efficiency and trade-offs.
LoRA's scaling factor dominates optimization by amplifying task signals without increasing drift, outperforming learning rate adjustments. The optimal alpha follows a sublinear square-root law with rank, revealing insufficient scaling in existing heuristics. Proposed LoRA-alpha restores principled s
GUI-SD uses on-policy self-distillation with privileged visual contexts and entropy-guided distillation for GUI grounding, outperforming GRPO methods in accuracy and efficiency.
Dirichlet-Guided Group Forecasting reduces time-series over-smoothing by modeling multi-modal predictive distributions with Dirichlet-guided sampling, improving accuracy, diversity, and dynamical consistency.
AtomWorld-Mem restores latent hidden dynamical states from atomistic snapshots via multi-scale memory to improve long-horizon kinetic Monte Carlo evolution and transfer across unseen alloys.
Dynamical Adapter Fusion derives optimal coefficients via PAC-Bayes and Taylor expansion to fuse task-specific adapters into one global adapter, achieving state-of-the-art class-incremental learning results.
Adam converges with high probability on generalized-smooth objectives under only second-moment stochastic gradients, matching a sharp δ^{-1/2} confidence dependence and yielding expectation rates for p<1.
MineEvolve converts Minecraft execution feedback into structured skills and remedies via Monitor, Inducer, Curator, and Adaptor, improving long-horizon agent performance across planners.