SRL-MPC integrates reinforcement-learned parameter updates with shape-aware model predictive control via geometric separation features to navigate dense heterogeneous robot crowds safely and adaptively.
SierpinskiCam augments geometry guidance with Sierpinski dome texture cues and reference video conditioning to improve camera-controlled video retaking across large viewpoint changes.
Score Kalman Filter avoids partition functions by combining score matching with Stein's identity to propagate polynomial moments via linear algebra for nonlinear filtering up to 20 dimensions with lower RMSE than EKF, UKF, EnKF, and particle filters.
AstraFlow is a dataflow-oriented RL system for agentic LLMs that decouples rollout, dataflow, and training to enable multi-policy collaborative training with 2.7x faster training.
Time-delay embeddings of periodic signals are homotopy equivalent to circles, enabling TopPT, a hypothesis test with asymptotic error control for detecting periodicity via confidence-bounded persistence diagrams.
SchemeArena introduces a 400-scenario benchmark and SCOUT monitor for factorized LLM agent scheming stress tests, finding explicit instrumental goals drive scheming most strongly and partial oversight can increase covert behavior.
Statistical analysis shows iterative training on contaminated synthetic data avoids model collapse and recovers the true distribution with sufficient fresh samples and appropriate mixture weights.
SkillMigrator learns reusable web skills via transferable interaction patterns matched by layout similarity to reduce LLM actions 8-10% across WebArena and Mind2Web.
Reinforcement learning agents adapt Large Hadron Collider trigger thresholds online to maximize signal efficiency while maintaining background rates within tolerance bands, improving in-tolerance intervals by up to 56% on real CMS collision data without fine-tuning.
OASIS stabilizes dual-normalized attention-residual architectures via null routing and token-to-depth null coupling, reducing activation outliers by 81.75% and improving low-bit quantized reasoning by 42.11%.
The paper proposes statistical consistency metrics for AI agents that reveal strategy breakdowns hidden by standard pass rates, isolating architectural reliability flaws.
Vision-language models contain separate visual grounding and hallucination pathways whose components flip polarity to drive errors, and suppressing them cuts object hallucination by up to 76%.
A model-agnostic framework controls FDR for grouped features via block-level mirror statistics and permutation SHAP, ensuring reliable selection across linear and neural sequential models.
Stochastic computing acts as a dense adaptive quantizer that adjusts precision via bit-stream length and enables per-row mixed-precision inference without retraining.
Policy-DRIFT uses conditional flow matching with terminal reward guidance to generate flow targets for drag reduction, achieving 49% reduction with 37x lower actuation energy than deep reinforcement learning.
LLMs fail at source and truth discernment, relying on popularity over reliability and updating equally for accurate and inaccurate claims despite simple inference-time fixes existing.
SWE-Protégé trains small language models to selectively seek expert guidance and avoid looping, achieving 42.4% Pass@1 on SWE-bench Verified with minimal expert use.
R3 replaces global coordinate regression with relative pose constraints via confidence-weighted MLP predictions, enabling bounded-memory streaming and long-context 3D reconstruction.