DPIAgent divides reproduction test generation into isolated diagnosis and test phases with structured handoffs, achieving up to 86.17% success on SWT-Bench Verified.
P2T uses reference patches as privileged supervision to curate shorter, grounded agent trajectories via bi-objective optimization, improving SWE-bench Pass@1 by up to 10.8 points with ~15% lower inference cost.
XL-DocBench introduces a human-verified benchmark for extra-long professional document understanding spanning thousands of pages with multi-page evidence, showing current systems still struggle with long-context structured reasoning.
DocAtlas treats long-document understanding as a mutable-state interaction process via a document harness with search, memory, and review tools, reaching 71.4% on MMLongBench-Doc and boosting a 4B VLM to 63.7% via reinforcement learning.
Model-guidance replaces classifier-free guidance by training on condition posterior probabilities, doubling inference speed and achieving 1.34 FID on ImageNet 256.
Memory Grafting uses frozen hidden states from a grafting model as conditional n-gram memory for language models, improving average benchmarks to 53.86 versus 52.43 for vanilla Engram at 2.8B scale with minimal overhead.
G2PO transforms agent trajectories into state-transition graphs to reduce variance and improve credit assignment, outperforming GRPO by up to 22.2% on long-horizon benchmarks.
Standardized item-level benchmark releases should become AI evaluation infrastructure because aggregate scores obscure validity failures; OpenEval archives 10M responses to enable auditability and recover benchmark validity evidence.
OneVision-Encoder applies codec-aligned sparsity to video, processing only high-entropy regions to outperform dense backbones with fewer tokens. It achieves 4.1% higher video accuracy than Qwen3-ViT across 16 benchmarks.
LCVN introduces a language-conditioned navigation benchmark and compares diffusion-based latent imagination against unified autoregressive prediction for embodied agents. Latent imagination yields more temporally coherent rollouts, while unified prediction generalizes better to unseen environments.