Alem benchmarks open-ended multi-agent coordination for language agents, showing frontier LLMs average ~6% returns and individual competence does not imply coordination competence.
Analytical Bias Correction fixes O(1/n) minibatch centroid bias in drifting models via a closed-form plug-in, reducing it to O(1/n²) with negligible overhead and improving CIFAR-10 FID.
EyeVLM benchmarks vision-language models on gaze following and social gaze prediction, finding they lack precise gaze understanding despite training improvements.