Event Cascade Pruning uses event-camera motion cues to prune video tokens in first-person spatial reasoning, improving accuracy by 1.31 points with 80% fewer tokens and 1.89x speedup.
A unified framework decomposes LLM alignment dynamics into competing rebound and driving forces, explaining reversal and faster re-alignment via rehearsal priming.
UniGraphLM proposes a unified graph language model that multi-domain multi-task aligns GNN representations to LLMs via adaptive alignment for cross-domain generalization.
MM-IssueLoc benchmarks multimodal repository-level issue localization using visual evidence across 652 instances, showing current systems achieve under 39% file accuracy and text-only scores do not transfer.
Deriving an alignment imprint from LLM preference tuning, LAPD detects AI-generated text with 45.82% relative gains over baselines via statistically guaranteed preference discrepancy.
DriveDreamer-Policy unifies depth generation, video prediction, and motion planning via geometry-aware world representations, achieving 89.2 PDMS on Navsim v1 and 88.7 EPDMS on v2.
MyoChallenge 2025 benchmarks musculoskeletal sports control via simulated table tennis and soccer tasks, advancing agile motor algorithms across 70 teams.
VETime unifies temporal and visual modalities via fine-grained alignment and dynamic fusion for zero-shot time-series anomaly detection, outperforming state-of-the-art models with lower overhead.
ECHO-2 is a distributed RL framework that overlaps rollout generation, dissemination, and training with bounded policy staleness to improve cost efficiency while preserving rewards.
A visual-native harness with an image bank and on-policy data evolution improves multimodal deep search agents, raising Qwen3-VL-8B to 39.0% average and surpassing Gemini-2.5 Pro.
Balanced Fine-Tuning uses dual-scale token and sequence reweighting targeting dense epistemic uncertainty to align LLMs with biomedical knowledge, improving reasoning and sparse-reward RL over standard fine-tuning.
Memory Grafting uses frozen hidden states from a grafting model as conditional n-gram memory for language models, improving average benchmarks to 53.86 versus 52.43 for vanilla Engram at 2.8B scale with minimal overhead.
DiffICL frames tabular synthesis as in-context learning using pretrained structural priors to avoid memorization, improving both quality and privacy in small-data settings.
CORAL introduces a structure-aware benchmark for local and whole-brain neuron reconstruction in light microscopy, revealing current methods need improvement for complete brain-wide tracing.
GeCO replaces fixed-schedule flow matching with time-unconditional optimization that adaptively allocates inference compute and uses field norms as training-free OOD detectors.
FacePhys is a memory-efficient rPPG algorithm using temporal-spatial state space duality that cuts error by 49% with a 3.6 MB footprint and 9.46 ms latency.
HACRL enables heterogeneous agents to share verified rollouts during collaborative on-policy training and execute independently at inference, with HACPO improving all agents by 3.6% over baselines at half the rollout cost.
LLM-updated agent memories degrade with repeated consolidation, often dropping below baselines; retaining raw episodic traces doubles accuracy versus forced consolidation.
ElegantVLA accelerates vision-language-action models via adaptive compute scheduling, achieving up to 3.77x speedup and doubling control frequency without retraining.
DiffScore evaluates text with masked diffusion models using bidirectional context to eliminate positional bias and decompose quality into fluency and faithfulness, outperforming autoregressive baselines.
SAM 3D Animal is a promptable framework using the SMAL+ model and Herd3D dataset to reconstruct multiple 3D animals from single images with keypoint and mask prompts, achieving state-of-the-art results.
cIPO aligns text-to-video diffusion by deriving implicit preferences from reconstruction errors and concentrating optimization on high-error temporal segments to fix sparse artifacts.
TELEVAL benchmarks Chinese spoken language models on interactional audio-conditioned dialogue, finding competitive semantic accuracy but degraded interactional performance under acoustic variability and a recurring caption trap failure.
LatentUM unifies modalities in a shared latent space to enable efficient interleaved cross-modal reasoning and generation, achieving state-of-the-art visual planning and self-reflective generation results.
SwiftVLM introduces cross-layer token bypass to preserve visual tokens across pruning stages, enabling training-free vision-language model acceleration with superior accuracy-efficiency trade-offs.
FML-bench isolates agent strategy from infrastructure across 18 ML tasks, finding greedy hill-climbing nearly matches tree search, while adaptive exploration switching outperforms fixed strategies.
TaskGround grounds full household scenes into task-relevant slices to infer executable task structures, improving compact open-weight models' success rates by large margins over direct prompting while cutting token costs up to 18x.
IRPO applies GRPO post-training to image restoration via selective hard-sample data and multi-component rewards, improving in-domain accuracy by 0.93 dB and OOD generalization by 3.43 dB.
CellMSA improves single-cell representation learning by modeling cross-batch and cross-cell-type context via MSA-inspired gene-pair representations, outperforming existing methods across benchmarks.
CausalMix frames data mixture optimization as causal inference to dynamically estimate optimal mixtures via conditional average treatment effects, improving LLM performance without retraining proxy models.
DashAttention uses adaptive α-entmax to select variable key-value blocks per query, enabling fully differentiable hierarchical sparse attention that matches full-attention accuracy at 75% sparsity with faster inference than FlashAttention-3.
UniVLR unifies text and visual reasoning into a shared visual workspace, using compressed visual latent tokens to outperform prior methods with fewer reasoning tokens.
AEvo formulates agentic evolution as an interactive environment where a meta-agent edits the evolution procedure to steer long-horizon search, achieving up to 26% relative improvement over baselines.
SATR constrains gradient-free RSNN updates via signal-adaptive KL trust regions, improving stability and matching PPO-LSTM returns on continuous control benchmarks.
MemSkill learns and evolves reusable memory skills for extracting and revising agent memories via selection, execution, and design loops, improving long-context task performance.
StructEvo uses structure-aware reinforcement learning with delta-structure fusion and hierarchical actions to outperform state-of-the-art protein directed evolution methods by up to 16.3%.
ScaleCUA scales computer use agents via verifiable synthetic tasks and efficient online RL, achieving new open-source state-of-the-art results on OSWorld and ScienceBoard.
MLS-Bench evaluates AI agents on inventing scalable ML methods across 140 tasks, finding current systems fail to reliably surpass human-designed approaches due to insufficient scientific validation insight.
Jet-Long dynamically rescales RoPE via bifocal local and long-range windows with an analytic length-aware schedule to extend LLM contexts without tuning, outperforming baselines on RULER, HELMET-RAG, and perplexity while retaining near-FlashAttention-3 throughput.