SimplexUQ benchmarks conformal wrappers on simplex-valued predictions, showing global calibration can hide severe under-coverage and no wrapper universally dominates across tasks and stratification maps.
This paper formalizes test-time scaling as an optimizable multi-LLM collaboration graph and proposes Agent-REINFORCE to search compute-optimal architectures under budget constraints. Experiments show it outperforms baselines in efficiency and finds graphs balancing accuracy with latency.