Symmetry in variational inference forces approximate minimizers to recover target statistics under misspecification, unifying prior results and yielding new directional guarantees.
A time-sensitive testing-by-betting framework favors early rejection via time-weighted rewards, yielding Bellman-optimal e-processes and an exponential-decay-optimal criterion recovering classical growth-rate optimality at large scales.
TokenSwap benchmarks and reduces MLLMs' modality gap by interleaving visual tokens with text, finding reasoning models have smaller gaps and training with TokenSwap mitigates it.
Linear probes on LLM residual streams identify a shared preference vector tracking pairwise choices across personas, with cross-persona transfer and causal steering.
Alem benchmarks open-ended multi-agent coordination for language agents, showing frontier LLMs average ~6% returns and individual competence does not imply coordination competence.
Value-filtered decoding selectively steers LLM generation using a value-based safety criterion with explicit false-intervention bounds, improving safety-utility trade-offs over baselines.
Deep MARL scales opinion dynamics to 1000 agents, finding high conformity in large networks reduces accuracy and promotes dishonesty, unlike small groups, revealing a mismatch with modern media.
In Bayesian multi-armed bandits, pretraining on expert data tightens regret bounds by mutual information with the optimal action, while an information-directed rule selects data sources maximizing immediate information gain, and trust inference safeguards against ineffective or compromised experts.
Regularized Newton training of overparameterized neural networks converges to a deterministic NNTK limit with exponentially fast uniform convergence across all frequencies, avoiding gradient descent's spectral bias.
Mean-field transformers exhibit rapid token distribution concentration onto projection-driven limits with explicit Wasserstein bounds scaling in inverse temperature β and time t.
Metropolis-adjusted Langevin correctors using score-based acceptance probabilities and a two-coin Bernoulli factory reduce diffusion model sampling bias and improve FID.
A parameter-efficient plugin extends frozen 10-second ECG foundation models to long, variable-length recordings via compatible long-sequence processing and semantically informed temporal modeling, outperforming sliding-window and pooling baselines.
A Bayes-assisted framework adaptively builds confidence sequences via predictive expected log-growth to achieve asymptotic log-optimality and narrower widths.
Fine-tuning LLMs on documents that flag claims as false makes them believe those claims, with belief rates jumping from 2.5% to 88.6%, though local negation phrasing largely prevents it.
A mean-field framework formulates inference-time diffusion control via weighted interacting particles to target distribution-level rewards with theoretical guarantees.
ARQ introduces a question generator that produces transferable intermediate stepping stones, improving reasoning LLM performance via fine-tuning on synthetic data.
Looped reasoning models converge to cyclic fixed points that stabilize attention and repeat feedforward inference stages iteratively, with recurrence size and normalization affecting stability.
MedMisBench reveals LLM medical accuracy collapses from 71% to 38% under misleading context, exposing a critical evaluation blind spot around epistemic resilience.
FedRepRAG is a federated RAG framework that exchanges only compact latent representations across clients to reduce inference overhead, outperforming local retrieval baselines on decentralized VQA and QA tasks.