Symmetry in variational inference forces approximate minimizers to recover target statistics under misspecification, unifying prior results and yielding new directional guarantees.
A time-sensitive testing-by-betting framework favors early rejection via time-weighted rewards, yielding Bellman-optimal e-processes and an exponential-decay-optimal criterion recovering classical growth-rate optimality at large scales.
TokenSwap benchmarks and reduces MLLMs' modality gap by interleaving visual tokens with text, finding reasoning models have smaller gaps and training with TokenSwap mitigates it.
Linear probes on LLM residual streams identify a shared preference vector tracking pairwise choices across personas, with cross-persona transfer and causal steering.
Alem benchmarks open-ended multi-agent coordination for language agents, showing frontier LLMs average ~6% returns and individual competence does not imply coordination competence.
Value-filtered decoding selectively steers LLM generation using a value-based safety criterion with explicit false-intervention bounds, improving safety-utility trade-offs over baselines.
Deep MARL scales opinion dynamics to 1000 agents, finding high conformity in large networks reduces accuracy and promotes dishonesty, unlike small groups, revealing a mismatch with modern media.
In Bayesian multi-armed bandits, pretraining on expert data tightens regret bounds by mutual information with the optimal action, while an information-directed rule selects data sources maximizing immediate information gain, and trust inference safeguards against ineffective or compromised experts.
Regularized Newton training of overparameterized neural networks converges to a deterministic NNTK limit with exponentially fast uniform convergence across all frequencies, avoiding gradient descent's spectral bias.
Mean-field transformers exhibit rapid token distribution concentration onto projection-driven limits with explicit Wasserstein bounds scaling in inverse temperature β and time t.
Metropolis-adjusted Langevin correctors using score-based acceptance probabilities and a two-coin Bernoulli factory reduce diffusion model sampling bias and improve FID.
A parameter-efficient plugin extends frozen 10-second ECG foundation models to long, variable-length recordings via compatible long-sequence processing and semantically informed temporal modeling, outperforming sliding-window and pooling baselines.
A Bayes-assisted framework adaptively builds confidence sequences via predictive expected log-growth to achieve asymptotic log-optimality and narrower widths.
Fine-tuning LLMs on documents that flag claims as false makes them believe those claims, with belief rates jumping from 2.5% to 88.6%, though local negation phrasing largely prevents it.
A mean-field framework formulates inference-time diffusion control via weighted interacting particles to target distribution-level rewards with theoretical guarantees.
ARQ introduces a question generator that produces transferable intermediate stepping stones, improving reasoning LLM performance via fine-tuning on synthetic data.
Looped reasoning models converge to cyclic fixed points that stabilize attention and repeat feedforward inference stages iteratively, with recurrence size and normalization affecting stability.
MedMisBench reveals LLM medical accuracy collapses from 71% to 38% under misleading context, exposing a critical evaluation blind spot around epistemic resilience.
FedRepRAG is a federated RAG framework that exchanges only compact latent representations across clients to reduce inference overhead, outperforming local retrieval baselines on decentralized VQA and QA tasks.
Value-based agents trained on diverse reward functions implicitly encode world models, extractable via P-learning, with sufficient conditions for exact dynamics recovery and cross-goal generalization.
STARE analyzes token-level entropy dynamics under GRPO, identifies a credit assignment mismatch, and uses surprisal-guided advantage reweighting to stabilize policy entropy, improving AIME accuracy by 4-8%.
Multi-teacher distillation pretrains EEG foundation models using vision and time-series teachers via masked latent denoising, outperforming self-supervised methods with 75% less pretraining data.
Simple baselines match or beat sophisticated code evolution across mathematical bounds, agent scaffolds, and ML competitions, revealing evaluation flaws and underscoring that expert-designed search spaces matter more than search algorithms.
GridProbe scores frame evidence via frozen VLM answer-space probing and adaptive selection to reduce long-video attention costs with minimal accuracy loss. It matches monolithic baselines on Video-MME-v2 at 3.36x lower compute and Pareto-dominates baselines on LongVideoBench.
KV-compressibility is a learnable property, so KV-CAT trains transformers via masked KV slots to yield representations more amenable to post-hoc compression without sacrificing quality.
Instruct-Particulate predicts articulated 3D part segmentation and joint parameters from meshes and kinematic specifications, scaling training via vision-language labels to improve cross-category and AI-generated mesh generalization.
Hyperagents integrate editable task and meta agents to enable metacognitive self-modification, with DGM-H improving across domains and accumulating meta-level improvements.
Structured Linear Controlled Differential Equations are universal time-series generators that approximate induced path laws on compact latent sets, and Generative SLiCEs improve probabilistic forecasting and downstream task performance on irregular grids.
Coherent hierarchical multi-label learning to defer uses selective-exclusion contracts to eliminate taxonomic deferral incoherence in medical imaging via projection and belief propagation.
The paper introduces Itô maps as any-step stochastic flow maps for SDEs that enable efficient single-pass future state prediction, posterior sampling, and inference-time control.
StraTA introduces trajectory-level strategies into agentic reinforcement learning via hierarchical rollout training, improving long-horizon decision-making and reaching 93.1% on ALFWorld.
RelAgent is an LLM agent that builds SQL feature queries and selects predictive models for relational learning, yielding fast, interpretable predictions deployable via standard databases.
A GNN encodes LTL instructions as Boolean formula sequences conditioning a policy, improving zero-shot multi-event RL instruction following in complex environments.
Poisoning LLM pretraining requires only ~250 malicious documents regardless of dataset or model scale, revealing constant-cost backdoor injection risks for large models.
Skill cascading attacks distribute malicious objectives across benign skills to harm agent systems, and SkillCascade reliably induces such failures while evading per-skill defenses.
SOAR proposes regression-based LiDAR relocalization for UAVs using locality-preserving sliding-window attention and coordinate-independent initialization, achieving state-of-the-art accuracy on UAVLoc with a 40% higher success rate and over 10 meters lower mean error.
Exact posterior score estimation derives closed-form posterior scores for linear Gaussian inverse problems, enabling efficient training and sampling that outperforms baselines with far fewer evaluations.
UI traces from LLM web agents identify underlying models with 96% F1 via passive JavaScript tracking, though randomized delays only partially mitigate fingerprinting.
Code2World uses renderable code generation for GUI world modeling, achieving top next-UI prediction and boosting Android navigation success by up to 9.5%.
A PAC learning framework for concurrent stochastic games computes robust near-optimal Nash equilibria with polynomial sample complexity or certifies nonexistence.
LibriBrain100 provides over 100 hours of MEG speech-decoding data, showing deep within-subject recordings and broad multi-subject data improve noninvasive word classification.
PPAT combines unbiased LURE estimation with prediction-powered control variates and adaptive acquisition to reduce label variance, yielding valid confidence intervals with fewer labels.
PAIR-CI is a calibrated nonparametric conditional independence test for incomplete data that uses paired cross-validated imputation to cancel imputation error, controlling false positives near nominal levels and improving causal discovery accuracy over existing methods.
Synthetic noise in pretraining data causes LLM loss divergence with probability scaling by noise type, amount, and model size, exhibiting activation patterns distinct from high-learning-rate failures.
Common interventions suppress emergent misalignment only under standard evaluations, yet hidden contextual triggers still elicit worse misalignment resembling training conditions.
Standard crosscoders learn layer-localized features; fmxcoders use factorized weights and layer masking to recover cross-layer features, improving coherence and reconstruction across four LLMs.
Predictively-Oriented Kalman Filter (EKF-PrO) uses fast approximate updates to avoid overconfident filtering under model misspecification without hyperparameters.
StemBind introduces a shared-stem benchmark diagnosing MLLM abstract visual reasoning, finding a persistent rule-to-instance binding gap where models identify patterns but fail to apply them correctly.
Strategic attack selection via start and stop policies substantially lowers measured AI control safety without changing attack capability, reducing safety by up to 28 percentage points and yielding overly optimistic estimates.