Gen-Searcher trains a search-augmented image generation agent via supervised and reinforcement learning, yielding about 16-point gains on knowledge-intensive benchmarks.
M²E-UAV introduces the first onboard event-based benchmark for motion-on-motion tiny UAV detection, showing current methods fail under dense ego-motion and sparse targets.
AlphaQ allocates MoE quantization bits without calibration using heavy-tailed spectral analysis, outperforming calibration-based methods and achieving near full-precision accuracy at 3.5-bit average precision.
DiPO disentangles perplexity into exploration and exploitation subspaces to enable fine-grained trade-offs, improving LLM reasoning and function calling via stable perplexity-guided policy optimization.
NAVA proposes native audio-visual alignment with an Align-then-Fuse MMDiT architecture for joint audio-video generation, achieving superior synchronization, video quality, and timbre control with 6.3B parameters.
OpenSearch-VL introduces an open-source recipe training multimodal deep search agents via curated data, diverse tools, and multi-turn fatal-aware GRPO, achieving over 10-point benchmark gains comparable to proprietary models.
FlowMAS learns multi-agent workflow topologies via reward-guided GFlowNets with structure-aware exploration and information-guided evaluation to outperform baselines.
ARGUS introduces multi-view identity mosaic injection and counterfactual training to preserve subject identity across motion, viewpoint changes, and occlusions in video generation.
EDA adapts draft models to fine-tuned LLMs via lightweight private components, regenerated training data, and selective sampling, restoring speculative decoding performance at much lower cost than full retraining.
SOAR proposes regression-based LiDAR relocalization for UAVs using locality-preserving sliding-window attention and coordinate-independent initialization, achieving state-of-the-art accuracy on UAVLoc with a 40% higher success rate and over 10 meters lower mean error.
ReflectMT internalizes translation reflection into direct inference via two-stage reinforcement learning, outperforming multi-step reasoning models with 94% fewer tokens.
ViCO minimizes vision tokens via consistency training across compression ratios, cutting tokens up to 50% while preserving capabilities through semantic-based routing.
AlphaPareto uses LLM-guided multi-objective reinforcement learning to discover formulaic trading alphas that adapt to evolving pools and outperform competitors on real-world data.
Unify-Agent reframes image synthesis as an agent pipeline with search and recaptioning, improving generation of long-tail factual concepts via 143K curated trajectories.