Sinks and diagonal patterns serve as attention switches and anti-oversmoothing mechanisms, with sinks favored in pretrained transformers due to lower representation costs.
Hamiltonian Causal Models separate equations of motion from intervenable mechanisms and define causal effects as interventional path discrepancies, showing entropy production witnesses trajectory-level causal effects invisible to standard average treatment effects.
Sparse autoencoders face a rate-distortion-polysemanticity tradeoff where monosemanticity raises reconstruction cost and data co-occurrence drives polysemanticity.
A per-sample trust score combining global realism and attribute-wise faithfulness evaluates conditional generations under compositional shift without reference data, enabling filtering and ranking that improves biological imaging and vision benchmarks.
Finite-horizon multi-environment POMDP optimization is PSPACE-complete, and a new practical algorithm significantly outperforms prior methods on benchmarks.