Quotienting Markov chains by their peripheral invariant subspace separates persistent regime profiles from transient dynamics for stable policy evaluation.
Truncated signature inversion is reframed as learning conditional path distributions via signature-conditioned flow matching, with derived Bayes error baselines and validated reconstruction on real data.
AgentForesight introduces online trajectory auditing that predicts multi-agent failures during execution, with a 7B model outperforming GPT-4.1 and DeepSeek-V4-Pro by up to 19.9% with 3x lower step error.
CURB penalizes policy dependence on defection histories via total-variation reward shaping to eliminate collusive equilibria in repeated multi-agent games.
Existing plasticity diagnostics fail to predict trainability, but optimization readiness, combining gradient strength and reliability, lower-bounds optimization gain and predicts plasticity more reliably.
SEED reinterprets decoder-only transformers as implicit encoder-decoders to reuse deep representations for fast self-speculative drafting, achieving up to 2.7x speedup.
CLReg uses contrastive regularization to separate forget and retain representations, reducing entanglement and improving LLM unlearning without extra privacy risks.
SARL improves reasoning via label-free reinforcement learning that rewards reasoning topology over outcomes, outperforming supervised and preference-based methods on math and open-ended tasks with more stable training.
DIDR aligns one-step generators via trajectory-level diffusion reward propagation, avoiding fidelity loss to Pareto-dominate SDXL and surpass 50-step teachers in one step.
On-Policy Consistency Training improves LLM safety across sycophancy, jailbreaks, and safety awareness while avoiding the capability regressions of supervised fine-tuning.
LLMs exhibit confirmation bias by proposing confirming triples rather than falsifying hidden rules, reducing discovery rates, though prompting counterexample consideration improves success from 42% to 56%.
Multiclass PAC learning with bandit feedback is characterized by the new bandit DS dimension via pseudo-boxes, yielding sharp sample complexity scaling with total neighbors.
SkillMigrator learns reusable web skills via transferable interaction patterns matched by layout similarity to reduce LLM actions 8-10% across WebArena and Mind2Web.
CORVUS decouples file reads from observations via synchronized registries, cutting input tokens by 9-50% and reasoning cycles by up to 37% while preserving pass rates.
Gated attention represents attention matrices as hierarchical mixtures of experts and achieves polynomial sample complexity versus exponential for multi-head self-attention.
Residual Latent Action predicts visual feature dynamics via flow matching, outperforming diffusion world models with orders-of-magnitude faster inference and enabling offline robot learning from videos.
Token value inequality in reasoning traces enables identifying core versus redundant tokens via log probabilities, yielding 76% token reduction with preserved accuracy via selective compression.