A geometric framework defines concept frustration as contradictions from missing concepts and detects it in foundation model embeddings to align human and machine reasoning.
Sobolev-regularized MMD gradient flow penalizes witness function gradients to ensure global convergence without isoperimetric assumptions, applying to both sampling and generative modeling.
High-dimensional analysis of pretraining via PCA and linear probing derives exact errors versus representation size, showing compression helps with abundant unlabeled but scarce labeled data.
MedMisBench reveals LLM medical accuracy collapses from 71% to 38% under misleading context, exposing a critical evaluation blind spot around epistemic resilience.
Fisher Decorator refines flow policies via local transport maps and Fisher-metric anisotropic optimization to fix isotropic approximation errors in offline RL.
Trace corpora collapse into compact finite-state machines replaying held-out data at >=0.997 fitness, yielding state-context next-step prediction and 0.94 AUROC failure prediction for runtime monitoring.
Model spec midtraining teaches models their behavior spec before alignment, controlling how demonstration fine-tuning generalizes and reducing agentic misalignment substantially.
CodeScaler uses a reward model to scale code LLM training and inference without test cases, improving benchmarks by up to 14.64 points and cutting latency tenfold.
OpenMHC releases the largest open wearable health dataset with open-source foundation models and a unified benchmark across prediction, imputation, and forecasting tasks.
Standard generative sequence models suffer physical misgeneralization, where local trajectory errors propagate through physical measurements to shift aggregate distributions; a data deviation kernel predicts these shifts and guides mitigation.
A neuro-symbolic framework trains a 4B-parameter model via MCTS-curated preference data and two-stage post-training to learn SAT cubing heuristics matching top symbolic methods.
Safety alignment effects in autonomous security agents require system-level measurement of refusal, tool reliability, and evidence grounding rather than refusal rates alone, with uncensored Gemma models improving security task success but showing mixed, family-dependent effects.
A Hilbert-valued one-step estimator enables semiparametrically efficient inference and bootstrap-calibrated tests for kernel noise heterogeneity in additive noise models.
Researchers derive maximally scale-stable parameterizations for Mixture-of-Experts via dynamical mean-field theory, yielding robust learning-rate transfer and monotonic scaling gains across regimes.
TerminalWorld automatically builds terminal benchmarks from wild recordings, yielding 1,530 tasks where top agents achieve only 62.5% success with weak correlation to expert benchmarks.