PRPO incorporates column-permutation invariance into LLM post-training via label-preserving permutations and two-level advantage estimation, enabling an 8B model to match specialized tabular baselines and outperform 685B reasoning LLMs by up to 53%.
LoopRPT applies reinforcement pre-training to looped language models by assigning rewards to latent reasoning steps, improving per-step quality and accuracy-computation trade-offs.
ROME uses role-playing LLMs to generate questionnaire answers from posts, then routes them via mixture-of-experts to improve personality detection and mitigate label scarcity.
Reformulating neural operators in d+1 dimensions via auxiliary embedding evolution achieves lowest relative L2 error across benchmarks without brute-force scaling.
Attention-based sampling orders diffusion language model tokens by attention-matrix column sums to maximize likelihood and improve generation quality with greater parallelism.
RePO-VLA improves vision-language-action robustness by assigning roles to success, recovery, and failure trajectories, raising adversarial success from 20% to 75%.
BALTO applies balanced token-level credit assignment to mitigate LLM hallucinations by redistributing probability from unsupported to faithful content, outperforming response-level methods on faithfulness benchmarks.
TransmissiveGS disentangles reflective and transmissive components via dual Gaussian representations and residual-guided multi-view separation for transmissive scene reconstruction.
Direct product flow matching decouples radial and angular dynamics via constant-speed geodesic transport and hidden-state conditioning for state-of-the-art few-shot vision-language adaptation.
BAS-VLA calibrates frozen VLA actions via breaking-centered calibration and selective preservation gating to suppress stale-task drift and separate semantics. It achieves 98% clean success, 0% under target swaps, and 70% under style shifts versus 42%.
VVTRec uses visual and textual visibility enrichment with vision-language models to reduce artifacts and improve radio interferometric image reconstruction.
HA-HOI reconstructs physically plausible 4D human-object interactions from monocular video by anchoring object motion to human action and refining via physics simulation. It improves alignment, contact consistency, and simulation readiness over prior monocular reconstruction methods.
ProteinOPD balances multi-objective protein preferences via on-policy distillation from preference-specific teachers, preserving designability with 8x training speedup over RL methods.
SMTL replaces sequential reasoning with parallel evidence acquisition for efficient long-horizon agentic search, achieving state-of-the-art results on multiple benchmarks with far fewer reasoning steps.
PROBE uses edit-response probing to build pocket-specific site maps and EditManuals that guide multi-agent optimization of both affinity and druggability in structure-based drug design, achieving state-of-the-art on CrossDocked2020.
TwinRouterBench introduces static and live dynamic tracks to benchmark LLM routing at agent step-level using deterministic scoring and live execution on SWE-bench.
A visual-native harness with an image bank and on-policy data evolution improves multimodal deep search agents, raising Qwen3-VL-8B to 39.0% average and surpassing Gemini-2.5 Pro.
Flexible Context Parallelism adaptively reconfigures communication groups to eliminate load imbalance and redundant communication, achieving up to 1.46x training speedup over Megatron-LM and DeepSpeed.
AHPA adaptively selects hierarchical VAE feature priors via a timestep-conditioned router to match diffusion transformer alignment granularity to denoising needs, improving convergence without inference overhead.
MedHorizon benchmarks long medical video understanding via sparse evidence retrieval and multi-hop reasoning, with top models reaching only 41.1% accuracy.
TACO is a training-free framework that learns adaptive compression rules from terminal agent trajectories to filter noisy observations, improving accuracy by 1-4% and reducing token usage across benchmarks.
A directional bias certificate enables Ellipsoidal-MINUCB to safely exploit offline data in linear contextual bandits, reducing regret when low-bias directions align with historical coverage.
AnchorWorld improves egocentric world simulation via full-body interaction supervision and anchor-view customization with consistent spatio-temporal dynamics.
CodeScaler uses a reward model to scale code LLM training and inference without test cases, improving benchmarks by up to 14.64 points and cutting latency tenfold.
VisInteract introduces interactive text-to-visualization with imperfect queries via VisInteract-Bench and Vis-MCTS, boosting success by over 13% versus interactive baselines.
Edit-R2 uses reinforcement learning to reconstruct session intent and jointly optimize reasoning and generation for multi-turn image editing. It improves instruction following and consistency over accumulated constraints on the MICE-Bench benchmark.
FocusDepth uses spatially-aligned multi-scale prompt fusion to boost target-region depth accuracy and sharp boundaries while preserving global geometry, outperforming global baselines on FDE-Bench.
HOMIE unifies inter- and intra-subject video personalization via multimodal guidance and reference embeddings, achieving state-of-the-art human-object interaction fidelity.
LLM-GNN Co-Teaching replaces golden-teacher design with bidirectional pseudo-label exchange and trajectory-based preference optimization, boosting few-shot graph accuracy by up to 7.86%.
HumanoidArena benchmarks egocentric hierarchical whole-body learning via seven leg-critical tasks, finding policies solve diverse interactions but cross-tracker transfer remains fragile.
VisHarness trains a visual agent to orchestrate heterogeneous experts for multi-turn reasoning, achieving strong results on segmentation, detection, and counting tasks.
A framework generates population-aligned personas from social media via quality filtering, importance sampling, and task-specific adaptation, reducing bias in LLM social simulations.
SpecBlock drafts block-iterative trees with hidden-state path dependence and adaptive branching, cutting EAGLE-3 drafting cost by about half while boosting speedup 8, 19%.
SSR3D-LLM introduces latent spatial reasoning steps to refine 3D object rankings step-by-step, improving fine-grained grounding across benchmarks while preserving unified language tasks.
Seg3DParts uses segmentation-grounded generation with structured cross-part interaction to produce controllable, coherent part-level 3D meshes from single images while introducing the PartObjectNet dataset.
DirectUV generates UV textures via diffusion with surface-aware positional encoding that attains 3D coherence across seams and improves occluded regions.
A framework maps compute budgets to optimal Mixture-of-Experts architectures via joint FLOP, active, and total parameter constraints, yielding robust scaling laws across hundreds of models with widening near-optimal flexibility at scale.
V-CAST prunes video tokens via curvature-guided temporal budgets and dual-anchor spatial selection, achieving 98.6% original performance with 86.4% latency.
AdaPreLoRA unifies LoRA optimizers via invertible Jacobian surrogates and Adafactor preconditioners, yielding efficient, accurate low-rank updates with minimal memory overhead.
CoDMD adds a copula-aware relational regularizer to distribution matching distillation that improves few-step video generation, achieving 84.46/84.87 VBench scores at 4 steps with ~25× speedup.
EvoMemBench benchmarks LLM agent memory via self-evolving scope and content axes, finding no universal memory method and that long-context baselines remain competitive.
Automatic multi-agent systems consistently underperform single-agent chain-of-thought self-consistency despite up to 10x cost, revealing automated architectures suffer from bloat and misaligned complexity rather than true multi-agent benefits.
HDR integrates hierarchical tree-structured latents into causal video generation to enable coarse-to-fine multi-step visual reasoning with sparse attention, boosting reasoning success by 76% over streaming diffusion while running 54x faster than bidirectional diffusion.
ProRes progressively warms up deeper layer residuals to stabilize pretraining, accelerating convergence and improving generalization across model scales.
MemCoRe organizes agent memory as a compression hierarchy to recover evidence from progressively compressed factual knowledge, outperforming state-of-the-art memory baselines.
IRPO applies GRPO post-training to image restoration via selective hard-sample data and multi-component rewards, improving in-domain accuracy by 0.93 dB and OOD generalization by 3.43 dB.
NGDB-Zoo improves neural graph database training via operator-level scheduling and semantic augmentation, achieving 1.8, 6.8x throughput without I/O stalls.