TwinRouterBench introduces static and live dynamic tracks to benchmark LLM routing at agent step-level using deterministic scoring and live execution on SWE-bench.
ECHO-2 is a distributed RL framework that overlaps rollout generation, dissemination, and training with bounded policy staleness to improve cost efficiency while preserving rewards.
A visual-native harness with an image bank and on-policy data evolution improves multimodal deep search agents, raising Qwen3-VL-8B to 39.0% average and surpassing Gemini-2.5 Pro.
SSR3D-LLM introduces latent spatial reasoning steps to refine 3D object rankings step-by-step, improving fine-grained grounding across benchmarks while preserving unified language tasks.
Flux Attention dynamically routes layer-level attention between full and sparse modes via a lightweight router to accelerate LLM inference, achieving up to 2.8x prefill and 2.0x decode speedups with minimal training.