BehaviorBench evaluates personalized decision modeling using real-world wallet traces across belief and trade prediction tasks, showing personalization improves beliefs more than trades and reveals model failure modes.
Procedural Memory Distillation extracts cross-episode strategy patterns into reusable procedural memory, co-evolving with the policy to improve reasoning benchmarks by up to 13.6% over SDPO.
Variational Policy Distillation co-evolves an adaptive teacher and student via variational EM to extract dense token-level guidance from language feedback, outperforming RLVR and self-distillation baselines on reasoning and code tasks.
Trajectory self-distillation trains few-step diffusion language models to match full-step trajectories, mitigating factorization error to enable fast, high-quality parallel decoding.
Automatic multi-agent systems consistently underperform single-agent chain-of-thought self-consistency despite up to 10x cost, revealing automated architectures suffer from bloat and misaligned complexity rather than true multi-agent benefits.
OrchRM uses self-supervised orchestration-level reward modeling to train multi-agent orchestrators, cutting token usage by 10x and boosting accuracy up to 8%.