Hidden-target projection minimizes regret for online inventory optimization on general convex sets, improving dependence on common-demand probability to inverse square root with matching lower bound, plus polylogarithmic and adaptive dynamic guarantees.
RIG-BENCH evaluates reasoning-driven image generation across four cognitive domains, revealing that state-of-the-art models produce locally plausible but globally illogical outputs.
OpenWebRL enables open online RL for visual web agents, with a 4B model reaching 67% Online-Mind2Web and 64% DeepShop success using minimal initialization data.
ELSA3D introduces elastic semantic anchoring to unify 3D understanding and generation via scale-matched cross-modal routing, achieving state-of-the-art results with roughly half the FLOPs and latency.
ΔLPS proposes a discrete gradient-informed posterior sampler that enables parallel updates without continuous relaxations, outperforming discrete diffusion samplers and matching continuous solvers across inverse problems.
RAW-Dream disentangles world models from task data by using task-agnostic pre-trained dynamics and VLM rewards to fine-tune VLAs entirely in zero-shot imagination with verified rollouts.
GPFlow learns generalized Poisson process rate functions for variable-length protein generation, improving designability and recovering length distributions without fixed-length constraints.
SimpliHuMoN is a simple transformer that predicts human pose and trajectory together, achieving state-of-the-art results across standard benchmarks without task-specific modifications.
A clustering-based divergence method measures gaps between real and simulated user behaviors, finding large, family-dependent discrepancies reducible by combining complementary simulators.
Lean4Agent uses Lean4 to formally model and verify agent workflows, with verified workflows outperforming failing ones by 11.94% and LeanEvolve improving SWE performance by 7.47%.
CoMMa is a decentralized multi-agent oncology framework using partitioned data, specialized fine-tuning, and deterministic contribution-aware aggregation to enable interpretable, privacy-preserving decision support across diverse clinical settings.
SILSA generates high-resolution 3D shapes with sliding-window slice latents and slice-level topology supervision to improve fidelity, reduce tokens by over 70%, and lower inference time by 58.5%.
STARFlow2 unifies multimodal generation by vertically interleaving a pretrained vision-language model with an autoregressive normalizing flow under shared causal masking, enabling cache-friendly interleaved text-image generation with strong benchmark performance.
CoT-Guard, a 4B-parameter chain-of-thought monitor, detects hidden code-generation objectives via SFT and RL, outperforming larger models including GPT-5.
Tailor-Bench evaluates visual world models on rare physical interactions via regular, unconventional, and impossible scenarios, revealing long-tail performance gaps and superficial visual-pattern reliance.
LLM-updated agent memories degrade with repeated consolidation, often dropping below baselines; retaining raw episodic traces doubles accuracy versus forced consolidation.
AFANet uses lightweight GNNs to attribute multi-agent failures via interaction graphs, matching LLM baselines with far lower cost and enabling test-time adaptation.
For discrete-time finite-horizon mean-field games with state-independent transitions and weakly monotone rewards, anchored proximal gradient descent computes mean-field equilibria via monotone inclusions over occupation measures at an O(1/√T) rate without regularization or uniqueness.
Exact posterior score estimation derives closed-form posterior scores for linear Gaussian inverse problems, enabling efficient training and sampling that outperforms baselines with far fewer evaluations.
Concept entanglement in diffusion models forces a trade-off where robust unlearning of a target necessarily damages overlapping concepts proportionally to their overlap.
A primal mirror descent algorithm computes exact Wasserstein barycenters for discrete and continuous measures in Fisher-Rao geometry with convergence guarantees.
VisAnomReasoner is a parameter-efficient vision-language model that improves time-series anomaly detection precision and F1 by over 21 points via VisAnomBench fine-tuning with natural-language rationales.
CELM is the first clinical EEG-to-language foundation model that generates long-duration EEG clinical reports and outperforms existing methods across all settings.
StableHand estimates world-space dual-hand motion from egocentric video via quality-aware flow matching, cutting W-MPJPE by 20-25% over baselines on occluded benchmarks.
ReToken introduces one learnable retrieval token that selects sparse visual tokens from long contexts, improving vision-language models by up to 13.4 points on visual retrieval while fitting on a single GPU.
This paper proposes Laws of Reasoning (LoRe), a framework formalizing reasoning compute and accuracy laws, plus LoRe-Bench showing models lack compositionality; enforcing compute-law compositionality via finetuning improves reasoning performance.
LangFlow closes the continuous-discrete language-modeling gap via flow matching and a learnable noise schedule, matching discrete diffusion perplexity and exceeding autoregressive zero-shot results on four benchmarks.
OrchRM uses self-supervised orchestration-level reward modeling to train multi-agent orchestrators, cutting token usage by 10x and boosting accuracy up to 8%.
PU-HNO predicts high-fidelity indoor radio maps from low-fidelity ray-tracing outputs via a three-stage physics-unrolled cascade that captures reflection, diffraction, and scattering effects, outperforming training labels and baselines.
KV-PRM eliminates text re-encoding by scoring via pre-existing KV caches, reducing process reward modeling cost from quadratic to linear and cutting latency and FLOPs by orders of magnitude.