GCPO replaces individual rollout scoring with team-level credit assignment based on valid solution coverage, significantly improving reasoning accuracy and diversity over competitive RLVR methods.
DiagEval uses trajectory-conditioned diagnostic probes to disambiguate evaluator errors from software defects in GUI-agent evaluations, recovering over 45% of misattributed failures and improving accuracy substantially.
AutoREM is a tuning-free memory-augmented framework that automates robust optimization reformulation via experience memory and improves accuracy across models.
SR-MCR aligns multimodal reasoning via intrinsic process rewards and a critic-free GRPO objective, achieving 81.4% average accuracy across visual benchmarks.
RePO-VLA improves vision-language-action robustness by assigning roles to success, recovery, and failure trajectories, raising adversarial success from 20% to 75%.
M3-AD proposes a reflection-aware multimodal benchmark and RA-Monitor framework that improves industrial anomaly detection via learnable self-correction, outperforming several MLLMs.
TSQAgent uses collaborative agent roles and external analytical tools to automatically identify relevant time series quality dimensions and perform quantitative comparisons, substantially improving LLM assessment and downstream data selection.
CodecSplat encodes 2D Gaussian-generation features into compact entropy-coded bitstreams to achieve order-of-magnitude smaller feed-forward 3D Gaussian splatting representations at high PSNR with controllable rate-distortion tradeoffs.
TPC treats captions as partial constraints, aligning vision-language representations to a consensus semantic core while penalizing dependence on unsaid residuals, yielding robust zero-shot recognition and improved LVLM grounding.
World models are formalized as group actions to enforce compositional dynamics via identity, inverse, and composition consistency, improving structural metrics without harming visual quality.
Balanced Fine-Tuning uses dual-scale token and sequence reweighting targeting dense epistemic uncertainty to align LLMs with biomedical knowledge, improving reasoning and sparse-reward RL over standard fine-tuning.
SIGMA uses semantic feature differencing with instruction-guided spatial priors to generate manipulation masks from edited images, producing a 1.1M training set that improves detectors by +18.34% F1.
DrawingsDreamer is a unified LLM-driven sequence model that generates multi-view engineering SVG drawings with high geometric fidelity and cross-view alignment via hierarchical tokenization and progressive training.
VisHarness trains a visual agent to orchestrate heterogeneous experts for multi-turn reasoning, achieving strong results on segmentation, detection, and counting tasks.
Delta-Adapter extracts a semantic delta from single image pairs to train exemplar-based editors without paired examples, improving accuracy and generalization.
cIPO aligns text-to-video diffusion by deriving implicit preferences from reconstruction errors and concentrating optimization on high-error temporal segments to fix sparse artifacts.
V-CAST prunes video tokens via curvature-guided temporal budgets and dual-anchor spatial selection, achieving 98.6% original performance with 86.4% latency.
Code2World uses renderable code generation for GUI world modeling, achieving top next-UI prediction and boosting Android navigation success by up to 9.5%.
IRR-Drive uses adaptive multimodal text and BEV reflection to self-correct driving intentions before trajectory generation, achieving state-of-the-art NAVSIM results.
NASDAQ normalizes low-dimensional observations to balance dynamics prediction losses and couples value learning with short-term value and next-observation prediction, achieving strong sample efficiency and faster training across diverse domains.
LiBrA-Net predicts low-resolution bilateral affine grids fused via Lie-algebraic regularization for real-time 4K video dehazing, and introduces the UHV-4K benchmark.