Hidden-target projection minimizes regret for online inventory optimization on general convex sets, improving dependence on common-demand probability to inverse square root with matching lower bound, plus polylogarithmic and adaptive dynamic guarantees.
RIG-BENCH evaluates reasoning-driven image generation across four cognitive domains, revealing that state-of-the-art models produce locally plausible but globally illogical outputs.
OpenWebRL enables open online RL for visual web agents, with a 4B model reaching 67% Online-Mind2Web and 64% DeepShop success using minimal initialization data.
ELSA3D introduces elastic semantic anchoring to unify 3D understanding and generation via scale-matched cross-modal routing, achieving state-of-the-art results with roughly half the FLOPs and latency.
ΔLPS proposes a discrete gradient-informed posterior sampler that enables parallel updates without continuous relaxations, outperforming discrete diffusion samplers and matching continuous solvers across inverse problems.
RAW-Dream disentangles world models from task data by using task-agnostic pre-trained dynamics and VLM rewards to fine-tune VLAs entirely in zero-shot imagination with verified rollouts.
GPFlow learns generalized Poisson process rate functions for variable-length protein generation, improving designability and recovering length distributions without fixed-length constraints.
SimpliHuMoN is a simple transformer that predicts human pose and trajectory together, achieving state-of-the-art results across standard benchmarks without task-specific modifications.
A clustering-based divergence method measures gaps between real and simulated user behaviors, finding large, family-dependent discrepancies reducible by combining complementary simulators.
Lean4Agent uses Lean4 to formally model and verify agent workflows, with verified workflows outperforming failing ones by 11.94% and LeanEvolve improving SWE performance by 7.47%.
CoMMa is a decentralized multi-agent oncology framework using partitioned data, specialized fine-tuning, and deterministic contribution-aware aggregation to enable interpretable, privacy-preserving decision support across diverse clinical settings.
SILSA generates high-resolution 3D shapes with sliding-window slice latents and slice-level topology supervision to improve fidelity, reduce tokens by over 70%, and lower inference time by 58.5%.
STARFlow2 unifies multimodal generation by vertically interleaving a pretrained vision-language model with an autoregressive normalizing flow under shared causal masking, enabling cache-friendly interleaved text-image generation with strong benchmark performance.
CoT-Guard, a 4B-parameter chain-of-thought monitor, detects hidden code-generation objectives via SFT and RL, outperforming larger models including GPT-5.
Tailor-Bench evaluates visual world models on rare physical interactions via regular, unconventional, and impossible scenarios, revealing long-tail performance gaps and superficial visual-pattern reliance.
LLM-updated agent memories degrade with repeated consolidation, often dropping below baselines; retaining raw episodic traces doubles accuracy versus forced consolidation.
AFANet uses lightweight GNNs to attribute multi-agent failures via interaction graphs, matching LLM baselines with far lower cost and enabling test-time adaptation.
For discrete-time finite-horizon mean-field games with state-independent transitions and weakly monotone rewards, anchored proximal gradient descent computes mean-field equilibria via monotone inclusions over occupation measures at an O(1/√T) rate without regularization or uniqueness.