GCPO replaces individual rollout scoring with team-level credit assignment based on valid solution coverage, significantly improving reasoning accuracy and diversity over competitive RLVR methods.
M3-AD proposes a reflection-aware multimodal benchmark and RA-Monitor framework that improves industrial anomaly detection via learnable self-correction, outperforming several MLLMs.
TSQAgent uses collaborative agent roles and external analytical tools to automatically identify relevant time series quality dimensions and perform quantitative comparisons, substantially improving LLM assessment and downstream data selection.
CodecSplat encodes 2D Gaussian-generation features into compact entropy-coded bitstreams to achieve order-of-magnitude smaller feed-forward 3D Gaussian splatting representations at high PSNR with controllable rate-distortion tradeoffs.
World models are formalized as group actions to enforce compositional dynamics via identity, inverse, and composition consistency, improving structural metrics without harming visual quality.
RefineAny3D treats monocular 3D depth refinement as visual alignment via categorical vision-language action tokens, boosting detectors without numerical regression.
A projected-gradient algorithm for Gromov-Wasserstein transport uses verifiable inexact projections to guarantee convergence to stationary points with scalable reliability.
Delta-Adapter extracts a semantic delta from single image pairs to train exemplar-based editors without paired examples, improving accuracy and generalization.
TIGER-FG uses text-guided implicit fine-grained grounding and dual distillation to improve cropped-query e-commerce retrieval, boosting Recall@1 by up to 34.4 points without object detection.
IRR-Drive uses adaptive multimodal text and BEV reflection to self-correct driving intentions before trajectory generation, achieving state-of-the-art NAVSIM results.
NASDAQ normalizes low-dimensional observations to balance dynamics prediction losses and couples value learning with short-term value and next-observation prediction, achieving strong sample efficiency and faster training across diverse domains.
LiBrA-Net predicts low-resolution bilateral affine grids fused via Lie-algebraic regularization for real-time 4K video dehazing, and introduces the UHV-4K benchmark.