POLAR-Bench evaluates LLM agent privacy-utility trade-offs via adversarial third-party probing across 10 domains, finding frontier models block over 99% of protected attributes while smaller open-weight models leak over half.
Mixed-policy LLM reasoning gains stem from buggy baselines; fixing optimizer and loss bugs makes standard SFT-then-RL outperform them by up to 22 points.