Post-training on protein-folding data via discrete answers and continuous geometry improves structure prediction and broad reasoning across ten benchmarks.
SkillOpt treats agent skills as external state optimized via bounded text edits validated on held-out scores, improving accuracy up to 24.8 points with stable transfer.
ARIS is an open-source autonomous research harness using cross-model adversarial collaboration to coordinate ML workflows and verify experimental claims.
DataFlex unifies sample selection, mixture adjustment, and reweighting for LLMs via a modular LLaMA-Factory framework that improves MMLU and perplexity with faster runtimes.
OPUS defines optimizer-induced update-space data utility for dynamic LLM pre-training selection, outperforming full-scale baselines with minimal overhead.
ADWM estimates LLM agent performance offline via a latent diffusion world model that alternates step-by-step with the policy, avoiding online interaction errors.
IAMFlow is a training-free identity-aware memory framework that tracks persistent entities across prompts to generate consistent long narrative videos, achieving best benchmark scores and faster inference.
TACache decomposes rectified flow velocity errors into magnitude and direction components to skip steps and reconstruct velocities without extra evaluations, achieving up to 4.14x faster image and 2.11x faster video generation.
MemForest partitions agent memory into event trees and progressively merges redundant nodes to cut storage and retrieval costs while preserving nearly all performance.
WorldAct converts static generated 3D worlds into editable, interaction-ready scenes via multimodal decomposition and object reconstruction to enable manipulation and embodied tasks.
PermaVid disentangles video memory into RGB appearance and depth structure with edit-aware updates to maintain long-term consistency across modifications.
Memory-R2 proposes LoGo-GRPO to enable fair credit assignment for memory-augmented LLM agents across long multi-session horizons via local rerollouts and shared-parameter co-learning.
D²Quant improves sub-4-bit LLM weight-only quantization via dual-scale quantizers for down-projection matrices and deviation-aware LayerNorm correction, boosting accuracy without extra bit budget.
PhotoFlow uses a Director-Reviewer-Reflector agent for closed-loop camera search to generate language-conditioned virtual photographs in arbitrary 3D scenes, outperforming baselines on quality, alignment, and success rate.
DeformGen uses dynamics-based topological augmentation to generate diverse deformable object states and warp trajectories for improved manipulation policy learning.
GLACIER treats tandem mass spectrum prediction as graph object detection, outperforming prior state-of-the-art by up to 19.3% on retrieval accuracy with nearly 8-fold faster inference.
Parallel LLM reasoning wastes compute via global budgets; sample-specific predictions via LanBo and PreAda improve efficiency without sacrificing accuracy.
HandEdit provides a 200M-instance benchmark and dataset for transforming egocentric human hands into diverse dexterous robot embodiments via image editing. It evaluates 11 baselines across hand-only and hand-arm tracks with embodiment-aware metrics.
AstroVLBench evaluates VLMs across five astronomical modalities, finding accuracy depends on physical grounding and raw numerical data improves results, yet all models lag behind domain-specialized methods.
IndustryCode is a multi-domain, multi-language benchmark of 579 industrial coding sub-problems; Claude 3.5 Opus reaches 68.1% sub-problem and 42.5% main-problem accuracy.
BALTO applies balanced token-level credit assignment to mitigate LLM hallucinations by redistributing probability from unsupported to faithful content, outperforming response-level methods on faithfulness benchmarks.
PO-PDDL learns symbolic POMDPs from robot videos to enable belief-space planning under partial observability and stochasticity, outperforming prior methods with lower planning cost.
OFBD identifies background-biased representations as a cause of long-tailed degradation and proposes foreground-guided CutMix and background-guided feature rectification to improve accuracy and tail-class performance.
xHC expands Transformer hyper-connections beyond four streams via sparse updates and temporal augmentation, improving scaling efficiency. It boosts 18B MoE downstream scores by 4.0 points over mHC with lower compute and reduced memory traffic via xHC-Flash.
Fisher Decorator refines flow policies via local transport maps and Fisher-metric anisotropic optimization to fix isotropic approximation errors in offline RL.
FocusDepth uses spatially-aligned multi-scale prompt fusion to boost target-region depth accuracy and sharp boundaries while preserving global geometry, outperforming global baselines on FDE-Bench.
OPERA jointly optimizes restoration planning via reinforcement learning and tool execution via co-training to outperform existing methods on complex mixed degradations.
ViDiHand leverages pretrained video diffusion models to reconstruct 4D hand poses directly from full egocentric video without detectors, substantially outperforming prior methods on ARCTIC, HOT3D, and HOI4D.
A GDRO algorithm via flexible sample queries and a prediction-with-limited-advice game achieves high-probability error O(√(∑ m/r_j)/t) with consistent sample complexity O(m log m/ε²).
DiffICL frames tabular synthesis as in-context learning using pretrained structural priors to avoid memorization, improving both quality and privacy in small-data settings.
DIO-Agent frames IO2Code as evolutionary search guided by execution errors and a simplicity-biased mutation prior, outperforming baselines on IO2CodeBench.
Justitia schedules task-parallel LLM agents via memory-centric cost prediction and virtual-time fair queuing to improve efficiency while preserving fairness and worst-case delays.
Quantum neural networks follow spectral amplitude priority rather than frequency bias, enabling efficient high-frequency learning via large-amplitude components and outperforming classical networks.
PLANING decouples geometry and appearance via explicit primitives and neural Gaussians for fast, high-quality streaming monocular 3D reconstruction with reduced redundancy.
V-CAST prunes video tokens via curvature-guided temporal budgets and dual-anchor spatial selection, achieving 98.6% original performance with 86.4% latency.
GUITAR diagnoses GUI agent failures via state transition graphs to reveal 60.4% of failures concentrate in 20% of bottleneck states, improving success rates with targeted guidance.
LatentUM unifies modalities in a shared latent space to enable efficient interleaved cross-modal reasoning and generation, achieving state-of-the-art visual planning and self-reflective generation results.