Benchmark aggregation is modeled as a principal-agent game where welfare loss depends on item alignment, improvability, and variance; auditing OLMES reveals Pareto-inferior items under pro-worker welfare.
Standard pass@k scaling laws suffer statistical shortcomings, so a beta-binomial framework and dynamic sampling strategy more accurately predict rare LLM capabilities and risks from limited data.
RADS applies reachability analysis and constrained reinforcement learning to steer diffusion trajectories away from memorized outputs via caption embedding perturbations, improving diversity, quality, and alignment without altering the model.
Standardized item-level benchmark releases should become AI evaluation infrastructure because aggregate scores obscure validity failures; OpenEval archives 10M responses to enable auditability and recover benchmark validity evidence.
TherapyGym introduces CTRS-based fidelity and multi-label safety evaluation for therapy chatbots, with RL training raising expert-rated CBT adherence from 0.10 to 0.60.
PU-DPO treats unmentioned radiology findings as unlabeled rather than negative, using edited contrastive pairs to prevent omission noise from corrupting preference optimization and improving hidden finding recovery.