A Riemannian ODE framework shows adaptive optimizers dampen flat directions too conservatively, and LITE accelerates Muon and SOAP by boosting flat-direction updates, cutting LLM pre-training time.
LITMUS benchmarks LLM agent behavioral jailbreaks in real OS environments, revealing agents execute 40.64% of high-risk operations despite refusals and suffer pervasive execution hallucination.