Quotienting Markov chains by their peripheral invariant subspace separates persistent regime profiles from transient dynamics for stable policy evaluation.
Truncated signature inversion is reframed as learning conditional path distributions via signature-conditioned flow matching, with derived Bayes error baselines and validated reconstruction on real data.
AgentForesight introduces online trajectory auditing that predicts multi-agent failures during execution, with a 7B model outperforming GPT-4.1 and DeepSeek-V4-Pro by up to 19.9% with 3x lower step error.
CURB penalizes policy dependence on defection histories via total-variation reward shaping to eliminate collusive equilibria in repeated multi-agent games.
Existing plasticity diagnostics fail to predict trainability, but optimization readiness, combining gradient strength and reliability, lower-bounds optimization gain and predicts plasticity more reliably.
SEED reinterprets decoder-only transformers as implicit encoder-decoders to reuse deep representations for fast self-speculative drafting, achieving up to 2.7x speedup.
CLReg uses contrastive regularization to separate forget and retain representations, reducing entanglement and improving LLM unlearning without extra privacy risks.
SARL improves reasoning via label-free reinforcement learning that rewards reasoning topology over outcomes, outperforming supervised and preference-based methods on math and open-ended tasks with more stable training.
DIDR aligns one-step generators via trajectory-level diffusion reward propagation, avoiding fidelity loss to Pareto-dominate SDXL and surpass 50-step teachers in one step.
On-Policy Consistency Training improves LLM safety across sycophancy, jailbreaks, and safety awareness while avoiding the capability regressions of supervised fine-tuning.
LLMs exhibit confirmation bias by proposing confirming triples rather than falsifying hidden rules, reducing discovery rates, though prompting counterexample consideration improves success from 42% to 56%.
Multiclass PAC learning with bandit feedback is characterized by the new bandit DS dimension via pseudo-boxes, yielding sharp sample complexity scaling with total neighbors.
SkillMigrator learns reusable web skills via transferable interaction patterns matched by layout similarity to reduce LLM actions 8-10% across WebArena and Mind2Web.
CORVUS decouples file reads from observations via synchronized registries, cutting input tokens by 9-50% and reasoning cycles by up to 37% while preserving pass rates.
Gated attention represents attention matrices as hierarchical mixtures of experts and achieves polynomial sample complexity versus exponential for multi-head self-attention.
Residual Latent Action predicts visual feature dynamics via flow matching, outperforming diffusion world models with orders-of-magnitude faster inference and enabling offline robot learning from videos.
Token value inequality in reasoning traces enables identifying core versus redundant tokens via log probabilities, yielding 76% token reduction with preserved accuracy via selective compression.
CCVFM replaces isotropic noise with a coreset-derived Gaussian mixture source for hierarchical rectified flow, using a lightweight correction flow for residuals to achieve competitive few-step generation.
Multi-site PPG is an in-the-wild dataset of 350+ hours from earring, ring, watch, and necklace wearables, showing heart-rate errors vary substantially by body site.
Under misspecified preference oracles, online LLM alignment minimizes a worst-case objective that decomposes into loss plus a sensitivity penalty, with projected updates achieving near-optimal complexity.
Concave scalarized multi-objective RL suffers biased gradients that cause O(ε⁻⁴) sample complexity; multi-level Monte Carlo NPG achieves optimal O(ε⁻²).
StreamGaze introduces a benchmark for evaluating gaze-guided temporal and proactive reasoning in streaming videos, revealing large performance gaps between state-of-the-art MLLMs and humans.
Analytical Bias Correction fixes O(1/n) minibatch centroid bias in drifting models via a closed-form plug-in, reducing it to O(1/n²) with negligible overhead and improving CIFAR-10 FID.
D-BOS differentiates through k-step softmax-Bayes belief dynamics to shape multi-agent opponent beliefs, outperforming PPO and BBM in hidden-role games.
Orthrus unifies autoregressive and diffusion views in transformers to enable lossless parallel token generation with up to 7.8x speedup and O(1) memory overhead.
RLVR trains a 30B LLM buyer via verifiable economic rewards to negotiate, revealing four-phase strategic evolution and outperforming much larger frontier models in surplus extraction.
NASDAQ normalizes low-dimensional observations to balance dynamics prediction losses and couples value learning with short-term value and next-observation prediction, achieving strong sample efficiency and faster training across diverse domains.
This paper bounds the query complexity of multi-round local search on general graphs, proving deterministic upper and randomized lower bounds that extend grid results to arbitrary connected graphs.
MAPs introduces a mini amusement-park simulator benchmarking integrated business decision-making, finding experts outperform state-of-the-art agents by over 11x due to weaknesses in long-horizon planning, sample-efficient learning, and spatial reasoning.
MLS-Bench evaluates AI agents on inventing scalable ML methods across 140 tasks, finding current systems fail to reliably surpass human-designed approaches due to insufficient scientific validation insight.