A time-sensitive testing-by-betting framework favors early rejection via time-weighted rewards, yielding Bellman-optimal e-processes and an exponential-decay-optimal criterion recovering classical growth-rate optimality at large scales.
Language generation in the limit is recast as recall-precision trade-offs, showing that allowing infinitely many vanishing-frequency hallucinations can strictly increase recall when adversaries withhold target portions.
PGLD optimizes synthesis-aware stochastic DNA libraries via policy gradients to bypass synthesis cost limits, enabling million-sequence libraries for antibody exploration at low cost.
Targeted Full Conformal Prediction uses vision-language models to prune labels and scale full conformal image classification with stable coverage and modest overhead.
Linear probes on LLM residual streams identify a shared preference vector tracking pairwise choices across personas, with cross-persona transfer and causal steering.
Multiple grids per group improve 4-bit quantization by selecting better grids per group, consistently boosting accuracy over single-grid FP4 for weights and activations.
CLoSeR detects loop candidates via global descriptors to close loops in streaming reconstruction, reducing drift and producing consistent kilometer-scale geometry with SE(3) pose optimization.
GenRec separates reconstruction and generation via observation masks to preserve fidelity in visible regions while synthesizing plausible unobserved content.
Benchmarking thirty metrics on ten simulated complex systems shows only causal metrics reliably validate high-level explanations when testing unmapped-variable faithfulness, leading to the Causal Abstraction Error metric converging with thirty interventions.
Prediction-intervention games model leaders choosing predictors against followers intervening on covariates; stable-blanket predictors are provably optimal or near-optimal.
Divide et Calibra uses vector quantization to learn shared, region-specific multiclass calibration maps that improve local calibration without reducing latent dimensions.
Under local PŁ conditions, unique optimistic lower-level selection ensures hyper-gradient differentiability via pseudoinverses, yielding HG-MS with manifold-dependent convergence and strong LLM reweighting results.
Morph is a flexible-size generative model for 3D molecular design that uses unbalanced optimal transport to dynamically adapt molecular size, improving property steering and enabling out-of-distribution generation.
Active Flow Expansion uses verifier-guided active exploration to grow a flow model's generable set, yielding theoretical guarantees and superior out-of-distribution molecule and protein design.
ProofRank benchmarks LLM proof quality via conciseness, ease, simplicity, diversity, and adaptivity, revealing quality differences and trade-offs with correctness.
Njord is a probabilistic graph neural network that produces efficient ensemble ocean forecasts via single-pass sampling, achieving lowest average upper-ocean errors globally and in the Baltic Sea.
ANCRe learns residual connectivities from data to fix convergence gaps caused by fixed layouts, accelerating training of deep networks with under 1% overhead.
TIDES moves input dependence from step size to the state matrix in selective SSMs, preserving physical time steps and per-token expressivity for irregular series, achieving state-of-the-art time-series results.
Anchor PCA finds shared low-rank directions across domains by trading variance for cross-domain agreement, yielding robust embeddings that generalize to unseen domains.
Discrete diffusion models learn data support before frequencies because reverse edits scale by validity first and coefficients second; absorbing diffusion prioritizes validity-improving moves over uniform diffusion's trichotomy.
Mixed-policy LLM reasoning gains stem from buggy baselines; fixing optimizer and loss bugs makes standard SFT-then-RL outperform them by up to 22 points.
AutoInject uses reinforcement learning with comparison-based rewards to learn adversarial suffixes that inject prompts into LLM agents, outperforming manual and optimization-based attacks on AgentDojo and Meta-SecAlign-70B.
FTerViT ternarizes all vision transformer weights and normalization parameters, achieving 82.43% ImageNet accuracy at 6.09MB and first ternary ViT deployment on ESP32-S3 microcontrollers.
Local sparsity in sparse autoencoder representations enables unsupervised LLM safety detection via masked activation analysis, achieving near-optimal detection using only 1-2% of neurons.
DéjàView loops a single transformer block for iterative multi-view 3D reconstruction, matching larger feed-forward models with far fewer parameters while treating refinement steps as an inference-time compute knob.
INEUS is a meshfree iterative neural solver for high-dimensional PIDEs that replaces nonlocal jump integrals with single-jump sampling and solves recursive regression problems with accurate scalable results.
Merge-Adversarial Training embeds durable text watermarks into open-source LLM weights via adversarial distillation, maintaining high detection rates after model merging while preserving capabilities.
A soft forward-backward algorithm learns stochastic policies from unlabeled data to optimize arbitrary differentiable occupancy utilities directly at test time via zero-order search over compact embeddings.
NoPo4D is the first feed-forward 4D Gaussian system that reconstructs dynamic scenes from unposed multi-view videos, outperforming baselines and optimization methods at much faster speeds.
Online discrete diffusion adaptation for molecular optimization finds acquisition, reward shaping, and debiasing complementarily boost reward, with replay and validity control stabilizing exploration to outperform offline and search baselines.
A binomial multibit LLM watermark directly encodes every payload bit at each token via a stateful encoder, outperforming baselines on large payloads with high robustness and proposing per-bit confidence scoring.