PRPO incorporates column-permutation invariance into LLM post-training via label-preserving permutations and two-level advantage estimation, enabling an 8B model to match specialized tabular baselines and outperform 685B reasoning LLMs by up to 53%.
LoopRPT applies reinforcement pre-training to looped language models by assigning rewards to latent reasoning steps, improving per-step quality and accuracy-computation trade-offs.
ROME uses role-playing LLMs to generate questionnaire answers from posts, then routes them via mixture-of-experts to improve personality detection and mitigate label scarcity.
Reformulating neural operators in d+1 dimensions via auxiliary embedding evolution achieves lowest relative L2 error across benchmarks without brute-force scaling.
Attention-based sampling orders diffusion language model tokens by attention-matrix column sums to maximize likelihood and improve generation quality with greater parallelism.
RePO-VLA improves vision-language-action robustness by assigning roles to success, recovery, and failure trajectories, raising adversarial success from 20% to 75%.
BALTO applies balanced token-level credit assignment to mitigate LLM hallucinations by redistributing probability from unsupported to faithful content, outperforming response-level methods on faithfulness benchmarks.
TransmissiveGS disentangles reflective and transmissive components via dual Gaussian representations and residual-guided multi-view separation for transmissive scene reconstruction.
Direct product flow matching decouples radial and angular dynamics via constant-speed geodesic transport and hidden-state conditioning for state-of-the-art few-shot vision-language adaptation.
BAS-VLA calibrates frozen VLA actions via breaking-centered calibration and selective preservation gating to suppress stale-task drift and separate semantics. It achieves 98% clean success, 0% under target swaps, and 70% under style shifts versus 42%.
VVTRec uses visual and textual visibility enrichment with vision-language models to reduce artifacts and improve radio interferometric image reconstruction.
HA-HOI reconstructs physically plausible 4D human-object interactions from monocular video by anchoring object motion to human action and refining via physics simulation. It improves alignment, contact consistency, and simulation readiness over prior monocular reconstruction methods.
ProteinOPD balances multi-objective protein preferences via on-policy distillation from preference-specific teachers, preserving designability with 8x training speedup over RL methods.
SMTL replaces sequential reasoning with parallel evidence acquisition for efficient long-horizon agentic search, achieving state-of-the-art results on multiple benchmarks with far fewer reasoning steps.
PROBE uses edit-response probing to build pocket-specific site maps and EditManuals that guide multi-agent optimization of both affinity and druggability in structure-based drug design, achieving state-of-the-art on CrossDocked2020.
TwinRouterBench introduces static and live dynamic tracks to benchmark LLM routing at agent step-level using deterministic scoring and live execution on SWE-bench.
A visual-native harness with an image bank and on-policy data evolution improves multimodal deep search agents, raising Qwen3-VL-8B to 39.0% average and surpassing Gemini-2.5 Pro.
Flexible Context Parallelism adaptively reconfigures communication groups to eliminate load imbalance and redundant communication, achieving up to 1.46x training speedup over Megatron-LM and DeepSpeed.
AHPA adaptively selects hierarchical VAE feature priors via a timestep-conditioned router to match diffusion transformer alignment granularity to denoising needs, improving convergence without inference overhead.
MedHorizon benchmarks long medical video understanding via sparse evidence retrieval and multi-hop reasoning, with top models reaching only 41.1% accuracy.
TACO is a training-free framework that learns adaptive compression rules from terminal agent trajectories to filter noisy observations, improving accuracy by 1-4% and reducing token usage across benchmarks.
A directional bias certificate enables Ellipsoidal-MINUCB to safely exploit offline data in linear contextual bandits, reducing regret when low-bias directions align with historical coverage.
AnchorWorld improves egocentric world simulation via full-body interaction supervision and anchor-view customization with consistent spatio-temporal dynamics.
CodeScaler uses a reward model to scale code LLM training and inference without test cases, improving benchmarks by up to 14.64 points and cutting latency tenfold.
VisInteract introduces interactive text-to-visualization with imperfect queries via VisInteract-Bench and Vis-MCTS, boosting success by over 13% versus interactive baselines.
Edit-R2 uses reinforcement learning to reconstruct session intent and jointly optimize reasoning and generation for multi-turn image editing. It improves instruction following and consistency over accumulated constraints on the MICE-Bench benchmark.
FocusDepth uses spatially-aligned multi-scale prompt fusion to boost target-region depth accuracy and sharp boundaries while preserving global geometry, outperforming global baselines on FDE-Bench.
HOMIE unifies inter- and intra-subject video personalization via multimodal guidance and reference embeddings, achieving state-of-the-art human-object interaction fidelity.
LLM-GNN Co-Teaching replaces golden-teacher design with bidirectional pseudo-label exchange and trajectory-based preference optimization, boosting few-shot graph accuracy by up to 7.86%.
HumanoidArena benchmarks egocentric hierarchical whole-body learning via seven leg-critical tasks, finding policies solve diverse interactions but cross-tracker transfer remains fragile.
VisHarness trains a visual agent to orchestrate heterogeneous experts for multi-turn reasoning, achieving strong results on segmentation, detection, and counting tasks.
A framework generates population-aligned personas from social media via quality filtering, importance sampling, and task-specific adaptation, reducing bias in LLM social simulations.
SpecBlock drafts block-iterative trees with hidden-state path dependence and adaptive branching, cutting EAGLE-3 drafting cost by about half while boosting speedup 8, 19%.