57%Worth a look?Worth a lookVote to see the scoreNeurIPS 2026Qualified HealthU California, San FranciscoWaymoLLM evaluation & benchmarksEfficient evaluation and error pattern discovery for blackbox AI systemsMaxim Rabinovich, Harvineet Singh, Aman SinhaSydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026– ReadersNo votes yet1/20 AI panelreviewers recommend itReaders and the AI panel: vote on this paper to see what they said.Worth readingNot for meOnly vote on papers you've read. Sign in with GitHub to vote.AI panel: 1 of 20 reviewers recommend itlenient 1/5medium 0/10strict 0/5
88%Must read?Must readVote to see the scoreNeurIPS 2026U California, San FranciscoCenter for Research in FoundatioUCSFAI oversight & deceptionAdaptive auditing of AI systems with anytime-valid guaranteesAn adaptive auditing framework using anytime-valid betting tests rigorously evaluates AI failure modes with as few as 20 observations and certifies global robustness upon passing stringent audits.Siyu Zhou, Patrick Vossler, Venkatesh Sivaraman, Yifan Mai and 1 moreSydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026– ReadersNo votes yet15/20 AI panelreviewers recommend itReaders and the AI panel: vote on this paper to see what they said.Worth readingNot for meOnly vote on papers you've read. Sign in with GitHub to vote.AI panel: 15 of 20 reviewers recommend itlenient 5/5medium 7/10strict 3/5