AI GameStore proposes evaluating general intelligence via scalable synthesis of human games, finding frontier vision-language models score under 10% of human averages on most generated games.
Temporal Backtracking Search improves video reasoning by searching over the temporal axis and restarting from verified prefixes rather than resampling from scratch, achieving 22.7% versus 0.7% best-of-N out-of-distribution.
A decentralized agent economy using auctions and economic selection emerges multi-step reasoning and outperforms monolithic baselines without centralized coordination.
Out-of-distribution shifts rotate active subspaces, misaligning dictionary explainers; a geometry-adaptive realignment using unlabeled OOD activations closes the faithfulness gap and restores causal interpretability without training.
A hierarchical sparse autoencoder architecture explicitly models semantic concept hierarchies, improving reconstruction, interpretability, and efficiency in language model representations.