BehaviorBench evaluates personalized decision modeling using real-world wallet traces across belief and trade prediction tasks, showing personalization improves beliefs more than trades and reveals model failure modes.
Procedural Memory Distillation extracts cross-episode strategy patterns into reusable procedural memory, co-evolving with the policy to improve reasoning benchmarks by up to 13.6% over SDPO.
Variational Policy Distillation co-evolves an adaptive teacher and student via variational EM to extract dense token-level guidance from language feedback, outperforming RLVR and self-distillation baselines on reasoning and code tasks.
TerraVis evaluates world-grounded visual consistency in generated images via MLLM workflows, correlating best with human judgments while revealing substantial failures in top models.
Automatic multi-agent systems consistently underperform single-agent chain-of-thought self-consistency despite up to 10x cost, revealing automated architectures suffer from bloat and misaligned complexity rather than true multi-agent benefits.
A Bayes-factor utility framework selects memory turns by evidence of latent preference changes and regulates access accordingly. It outperforms embedding retrieval on preference-intensive long-context dialogue tasks.
OrchRM uses self-supervised orchestration-level reward modeling to train multi-agent orchestrators, cutting token usage by 10x and boosting accuracy up to 8%.