Deep Double Q-learning trains two independent Q-functions to decouple selection and evaluation, reducing overestimation and outperforming Double DQN across 47 of 57 Atari games.
HiFloat4 enables stable FP4 LLM pretraining without stabilization stacks, achieving 1.55% relative loss versus 1.79% for MXFP4 and 2.00% for NVFP4 on Ascend NPUs.
This work formalizes nonlinear measure-to-measure regression and introduces two scalable transformer-based approaches for learning operators between probability distributions. The methods generalize to unseen measures in synthetic experiments, particle systems, and a large-scale colorectal cancer or
Sech perturbation kernels make calibration functions analytic, enabling polynomial regression to estimate second-order calibration error at the minimax optimal rate of tilde O(1/sqrt(n)). This yields the first finite-sample guarantee for second-order Platt scaling and a bucket-free calibration defin
HiFloat4 enables end-to-end FP4 reinforcement learning by fixing rollout activation underflow with Rollout-ResQ, cutting accuracy gaps to 1.1% versus BF16.
Laplacian Keyboard hierarchically combines Laplacian eigenvectors into a behavior library with a meta-policy, exceeding linear span limits for better zero-shot approximation and sample efficiency.
NextLat adds latent self-prediction to transformers, theoretically converging to belief states and empirically improving world modeling, reasoning, and inference speed.