Rep2Text recovers roughly half of tokens from single LLM representations via adapter-based decoding, showing sequence-length bottlenecks preserve semantics but reduce token recovery.
ReFPO adds explicit reflow regularization to flow matching policy gradients, stabilizing training and enabling high-fidelity one-step inference that matches multi-step performance across control tasks.
GeoSym Engine automates symbolically-verifiable geometric reasoning data synthesis, and models trained on GeoSym127K achieve large gains on diagram-dependent geometry benchmarks.
UniClawBench introduces a capability-driven benchmark evaluating proactive agents via 400 real-world tasks with live Docker evaluation and multi-turn feedback.
Humanoid robot co-design evolves control and morphology together via strategic exploration and meta-policy learning to enable true embodied intelligence.
TRL-Bench standardizes cross-paradigm evaluation of tabular encoders via shared representation-level probes, finding encoder quality is task-specific and best pipelines combine capability-matched specialists.
FlowSteer lets an agent design executable agentic workflows via a canvas environment that returns syntax-checked feedback for each edit, significantly outperforming baselines across twelve datasets.
Forced Deferral Attack uses adversarial image triggers to suppress weak-model confidence and force multimodal LLM cascades to route queries to strong models. It learns universal border triggers via temperature-flattened optimization, consistently increasing unintended strong-model usage across datas
Skill cascading attacks distribute malicious objectives across benign skills to harm agent systems, and SkillCascade reliably induces such failures while evading per-skill defenses.
BusterX introduces GenBuster-200K, GenBuster-Bench, and an MLLM baseline that detects AI-generated video via reasoning chains, outperforming leading models in accuracy and explanation quality.
FTC-Seg uses orthogonal prototype reconstruction and adaptive threshold calibration to break pseudo-label degradation cycles between imaging noise and long-tail class imbalance in semi-supervised semantic segmentation.
MineEvolve converts Minecraft execution feedback into structured skills and remedies via Monitor, Inducer, Curator, and Adaptor, improving long-horizon agent performance across planners.
LLM agents suffer intrinsic over-calling bias from an activation-independent call offset, which sparse autoencoders diagnose and steering corrects to improve accuracy.
Evo-Depth is a lightweight 0.9-billion-parameter vision-language-action model using implicit depth encoding from RGB to improve spatial manipulation without extra sensors. It achieves top benchmark performance with minimal GPU memory and highest inference speed among compared methods.