Predictive representation learning combined with high-capacity value approximation drives scalable multitask RL, with the simple model-free MR.Q outperforming world-model methods across continuous control tasks.
Bidirectional Information Flow enables continuous two-way communication in hierarchical Gaussian processes for Bayesian optimization, improving sample efficiency, training robustness, and modular subtask reuse while significantly outperforming unidirectional and vanilla methods.
Looped reasoning models converge to cyclic fixed points that stabilize attention and repeat feedforward inference stages iteratively, with recurrence size and normalization affecting stability.