D³PO fixes multi-objective RL via decomposed per-objective updates, late preference weighting, and diversity regularization to recover dense Pareto fronts with one policy.
This paper formalizes test-time scaling as an optimizable multi-LLM collaboration graph and proposes Agent-REINFORCE to search compute-optimal architectures under budget constraints. Experiments show it outperforms baselines in efficiency and finds graphs balancing accuracy with latency.