Language generation in the limit is recast as recall-precision trade-offs, showing that allowing infinitely many vanishing-frequency hallucinations can strictly increase recall when adversaries withhold target portions.
PGLD optimizes synthesis-aware stochastic DNA libraries via policy gradients to bypass synthesis cost limits, enabling million-sequence libraries for antibody exploration at low cost.
Targeted Full Conformal Prediction uses vision-language models to prune labels and scale full conformal image classification with stable coverage and modest overhead.
PaGeR adapts perspective 3D foundation models to panoramas to predict depth, normals, and sky masks in one pass, achieving state-of-the-art 360-degree geometry estimation.
Under local PŁ conditions, unique optimistic lower-level selection ensures hyper-gradient differentiability via pseudoinverses, yielding HG-MS with manifold-dependent convergence and strong LLM reweighting results.
A byte-level sequential Monte Carlo algorithm samples from composed language model ensembles, outperforming naive probability averaging across structured generation tasks.
Class Adaptive Conformal Training adaptively shapes class-conditional prediction sets via augmented Lagrangian optimization without distributional assumptions, yielding smaller sets with valid coverage.
EchoPrune treats redundant video tokens as temporal echoes and prunes them via query relevance and reconstruction error, letting VideoLLMs process up to 20x more frames for +8.6% accuracy and 5.6x faster prefilling.
Anchor PCA finds shared low-rank directions across domains by trading variance for cross-domain agreement, yielding robust embeddings that generalize to unseen domains.
Lumberjack improves differentially private random forests via heavy hitter pruning of deep trees, achieving state-of-the-art privacy-utility trade-offs.
A discrete-time stochastic control formulation yields value-driven transport policies that generate data via straight, fast, robust paths and support conditional generation and guidance.
Discrete diffusion models learn data support before frequencies because reverse edits scale by validity first and coefficients second; absorbing diffusion prioritizes validity-improving moves over uniform diffusion's trichotomy.