PAGER closes the semantic-execution gap for point-precise geometric GUI control via dependency-structured planning and pixel-level execution, achieving 4.1x higher task success than general baselines.
OFBD identifies background-biased representations as a cause of long-tailed degradation and proposes foreground-guided CutMix and background-guided feature rectification to improve accuracy and tail-class performance.
Balanced Fine-Tuning uses dual-scale token and sequence reweighting targeting dense epistemic uncertainty to align LLMs with biomedical knowledge, improving reasoning and sparse-reward RL over standard fine-tuning.
RankE co-evolves discrete text-to-image policy and decoder via alternating optimization to eliminate latent covariate shift, improving both FID and CLIP scores.
Tailor-Bench evaluates visual world models on rare physical interactions via regular, unconventional, and impossible scenarios, revealing long-tail performance gaps and superficial visual-pattern reliance.
WAV introduces a latent-space planning framework for vision-language-action models that predicts future states and evaluates trajectory values to enable efficient long-horizon decision-making.
RLVR training causes reasoning outputs to structurally converge on seen prompts, and Min-kNN Distance detects this collapse via black-box sampling to identify contamination.
PACE infers continuous single-cell dynamics via geometry-aware Riemannian transport, reducing reconstruction distances by 23.7% without paired cells or velocity supervision.
Manifold drift pushes flow preference optimization off the data manifold via terminal displacement normal components; ThermoDPO-weighted improves strict score and image metrics over FlowDPO.
R3 replaces global coordinate regression with relative pose constraints via confidence-weighted MLP predictions, enabling bounded-memory streaming and long-context 3D reconstruction.
LCVN introduces a language-conditioned navigation benchmark and compares diffusion-based latent imagination against unified autoregressive prediction for embodied agents. Latent imagination yields more temporally coherent rollouts, while unified prediction generalizes better to unseen environments.