VideoRLVR applies reinforcement learning with verifiable rewards to video diffusion models, improving rule-consistent visual reasoning and cutting training latency 40% via early-step optimization.
Masked-input regularization improves autoregressive pretraining over weight decay alone, and SoftQ scaling laws better capture data-constrained training than Chinchilla.
Multimodal LLMs suffer spurious cross-modality interference that distorts decisions, and a unified finetuning framework with perturbation augmentation and consistency regularization improves robustness and generalization.
OceanCBM is a concept bottleneck model that predicts ocean heat content through physically derived intermediate concepts, yielding consistent interpretable representations without sacrificing predictive skill.
Nonasymptotic total variation analysis of Riemannian flow matching bounds sampling error by discretization and learning terms via curvature-aware differential inequalities. Explicit polynomial iteration complexities follow on hyperspheres and SPD manifolds.
SkillsBench benchmarks agent skills across 87 tasks, finding curated skills boost pass rates by 16.6 points, with focused small bundles often outperforming larger ones.
Exact Gaussian moment matching propagates mean and covariance through residual networks with exact nonlinear layer formulas, cutting KL divergence errors by orders of magnitude versus approximate methods.
Linear self-attention transformers provably implement in-context policy-improvement via explicit constructions, with gradient flow converging exponentially to optimal RL update parameters under distribution richness conditions.
ModelLens learns a latent space over model-dataset-metric tuples from noisy leaderboard data to rank unseen models on unseen datasets without target evaluation, improving routing by up to 81%.
Rank-transformed dissimilarity profiles convert high-dimensional observations into robust class-wise rank profiles that encode moment differences and improve low-sample-size classification.