PRPO incorporates column-permutation invariance into LLM post-training via label-preserving permutations and two-level advantage estimation, enabling an 8B model to match specialized tabular baselines and outperform 685B reasoning LLMs by up to 53%.
LoopRPT applies reinforcement pre-training to looped language models by assigning rewards to latent reasoning steps, improving per-step quality and accuracy-computation trade-offs.
ROME uses role-playing LLMs to generate questionnaire answers from posts, then routes them via mixture-of-experts to improve personality detection and mitigate label scarcity.
Reformulating neural operators in d+1 dimensions via auxiliary embedding evolution achieves lowest relative L2 error across benchmarks without brute-force scaling.
Attention-based sampling orders diffusion language model tokens by attention-matrix column sums to maximize likelihood and improve generation quality with greater parallelism.
RePO-VLA improves vision-language-action robustness by assigning roles to success, recovery, and failure trajectories, raising adversarial success from 20% to 75%.
BALTO applies balanced token-level credit assignment to mitigate LLM hallucinations by redistributing probability from unsupported to faithful content, outperforming response-level methods on faithfulness benchmarks.
TransmissiveGS disentangles reflective and transmissive components via dual Gaussian representations and residual-guided multi-view separation for transmissive scene reconstruction.
Direct product flow matching decouples radial and angular dynamics via constant-speed geodesic transport and hidden-state conditioning for state-of-the-art few-shot vision-language adaptation.
BAS-VLA calibrates frozen VLA actions via breaking-centered calibration and selective preservation gating to suppress stale-task drift and separate semantics. It achieves 98% clean success, 0% under target swaps, and 70% under style shifts versus 42%.
VVTRec uses visual and textual visibility enrichment with vision-language models to reduce artifacts and improve radio interferometric image reconstruction.
HA-HOI reconstructs physically plausible 4D human-object interactions from monocular video by anchoring object motion to human action and refining via physics simulation. It improves alignment, contact consistency, and simulation readiness over prior monocular reconstruction methods.
ProteinOPD balances multi-objective protein preferences via on-policy distillation from preference-specific teachers, preserving designability with 8x training speedup over RL methods.