Gen-Searcher trains a search-augmented image generation agent via supervised and reinforcement learning, yielding about 16-point gains on knowledge-intensive benchmarks.
Equilibrium Matching learns implicit energy landscapes for optimization-based sampling, surpassing diffusion models with 1.90 FID on ImageNet 256x256 while supporting denoising, OOD detection, and composition.
Equilibrium Forcing removes noise conditioning from video diffusion to enable adaptive closed-loop inference that improves generation quality and consistency.
AutoSpec is a neural framework that discovers iterative spectral algorithms via self-supervised prediction of recurrence coefficients, yielding order-of-magnitude speedups over classical baselines.
Spectral Feedback iteratively selects protein tokens to re-mask and resample using sparse Fourier edit-set value functions, improving inverse-folding stability by up to 32.3% at test time.
AlphaQ allocates MoE quantization bits without calibration using heavy-tailed spectral analysis, outperforming calibration-based methods and achieving near full-precision accuracy at 3.5-bit average precision.
ICWBench reveals current LLMs fail at in-context watermarking, and self-distillation with reinforcement learning raises watermark detectability near perfect while preserving quality.
TTCD uses a long-window teacher to supervise a short-window student's fast weights via hidden-state discrepancy, allocating limited memory to future-relevant context and outperforming existing long-context methods with minimal architectural changes.
2FFS adaptively combines cheap biased heuristics and expensive accurate rollouts to identify best actions in stochastic minimax trees with fewer samples than baselines.
Animation2Code benchmarks video-to-code generation for web animations, showing state-of-the-art vision-language models struggle with temporal consistency despite high appearance fidelity.
Recon scores reasoning traces by action reconstruction fidelity to avoid post-hoc rationalization in user modeling, yielding up to 70% win rates over baselines across domains.
SCDBench benchmarks LLM smart-contract decompilers on 600 real contracts via semantic replay, finding even top models perfectly recover only 42 and same-model repair substantially helps.
AstraFlow is a dataflow-oriented RL system for agentic LLMs that decouples rollout, dataflow, and training to enable multi-policy collaborative training with 2.7x faster training.
MAGE uses block-diffusion's aligned all-[MASK] queries to select reusable sparse KV subsets, achieving near-lossless accuracy with up to 6.82x speedup at 128K context.
A clustering-based divergence method measures gaps between real and simulated user behaviors, finding large, family-dependent discrepancies reducible by combining complementary simulators.
BankerToolBench benchmarks AI agents on multi-hour investment banking workflows using expert rubrics, finding frontier models fail nearly half of criteria with zero client-ready outputs.
LaMo extracts self-supervised latent motion priors from unlabeled videos via motion drift loss and prior guidance, improving physical consistency in video diffusion without external supervision.
ReMind trains video diffusion transformers to use cache memory for evolving hidden states across interruptions via memory-oriented curricula and PM-RoPE, achieving best STEVO-Bench scores without catastrophic forgetting.
A dataset of multi-aspect human visual similarity judgments benchmarks vision-language models and yields the TPIPS metric, which aligns with human perception and enables text-guided image retrieval and generative evaluation.
Rosetta neuron populations grow sublinearly and become more selective and specialized as language and vision models scale, while non-Rosetta neurons stay less selective.
Poisoning LLM pretraining requires only ~250 malicious documents regardless of dataset or model scale, revealing constant-cost backdoor injection risks for large models.
DiscoLoop combines discrete embeddings and continuous hidden states in looping transformers to fix representational misalignment, enabling near-perfect multi-hop reasoning with faster training and stronger pretraining performance.
M²RNN introduces matrix-valued non-linear RNNs that scale via state expansion, achieving perfect state tracking and outperforming hybrid models with smaller states.
CFGRL links diffusion guidance to policy improvement, training via supervised learning to exceed dataset performance on offline RL without value functions.
Nearest-neighbor radii under mixing dependence converge almost surely with polynomial mixing and have sharp moment bounds scaling with local intrinsic dimension, remaining informative for high-dimensional dependent data.
Focal log-frequency loss balances spectral learning signals in flow matching, accelerating convergence by 40% and improving image fidelity without architectural changes.
A benchmark of likely public-domain Hollywood films tests multimodal models on narrative understanding, finding vision-language models near chance and audio-visual models below human performance.
NLAC trains LLM agents with a natural-language generative critic for off-policy learning, yielding richer feedback and more stable, data-efficient training than policy gradients in long-horizon tasks.
ACuRL enables autonomous continual learning for computer-use agents via curriculum reinforcement learning, yielding 3-29% gains without catastrophic forgetting or human data.