PRPO incorporates column-permutation invariance into LLM post-training via label-preserving permutations and two-level advantage estimation, enabling an 8B model to match specialized tabular baselines and outperform 685B reasoning LLMs by up to 53%.
ROME uses role-playing LLMs to generate questionnaire answers from posts, then routes them via mixture-of-experts to improve personality detection and mitigate label scarcity.
Reformulating neural operators in d+1 dimensions via auxiliary embedding evolution achieves lowest relative L2 error across benchmarks without brute-force scaling.
AdapToPASS is a bio-inspired spherical transformer that adaptively models contextual and geometric ambiguities to achieve robust panoramic semantic segmentation, outperforming state-of-the-art methods by up to 18.77% relative mIoU under unseen spherical transformations.
RotMoLE adds a rotation gate to MoE-LoRA experts that rotates rather than merely scaling selected experts, improving specialization and performance on multi-task and multilingual benchmarks.
Advanced optimizers weaken multi-task learning because instant gradients barely affect updates, so the APT framework with adaptive momentum and Muon direction preservation improves results across four datasets.
CodeScaler uses a reward model to scale code LLM training and inference without test cases, improving benchmarks by up to 14.64 points and cutting latency tenfold.
HumanoidArena benchmarks egocentric hierarchical whole-body learning via seven leg-critical tasks, finding policies solve diverse interactions but cross-tracker transfer remains fragile.
Next Forcing uses multi-chunk prediction to accelerate convergence 2.3x, boost high-frame-rate accuracy 93.1%, and double inference speed for world models.
Seg3DParts uses segmentation-grounded generation with structured cross-part interaction to produce controllable, coherent part-level 3D meshes from single images while introducing the PartObjectNet dataset.
XDecomposer learns prior-free multiphase X-ray diffraction decomposition as set prediction to identify constituent phases and proportions without candidate lists. It improves reconstruction accuracy and phase identification across simulated and experimental datasets.
IRR-Drive uses adaptive multimodal text and BEV reflection to self-correct driving intentions before trajectory generation, achieving state-of-the-art NAVSIM results.
EvoMemBench benchmarks LLM agent memory via self-evolving scope and content axes, finding no universal memory method and that long-context baselines remain competitive.
Automatic multi-agent systems consistently underperform single-agent chain-of-thought self-consistency despite up to 10x cost, revealing automated architectures suffer from bloat and misaligned complexity rather than true multi-agent benefits.
ColorConceptBench evaluates text-to-image models on probabilistic color associations for 1,281 implicit concepts, revealing substantial performance gaps and insensitivity to abstract semantics.
OmniTraffic introduces a controllable 3D traffic generation pipeline and benchmark with 8M VQA samples for spatio-temporal reasoning, revealing large model gaps and improved real-world performance via simulated fine-tuning.