Bayesian theory reveals softmax attention learns copy heads via a first-order data phase transition, unlike linear attention's second-order transition and crossover.
Croissant Baker generates validated Croissant metadata locally from dataset directories via modular handlers, achieving 97, 100% agreement with ground truth across 140+ datasets including MIMIC-IV.
Joint Self-Improvement uses a joint generative-predictive model and self-improving sampling to reduce distribution shift and efficiently generate optimized molecules under limited evaluation budgets.