57%Worth a look?Worth a lookVote to see the scoreNeurIPS 2026KhalifaMohamed bin Zayed University of Khalifa University of Science, TU Electronic Science and TechnolMedical imagingREFORM-3D: A Representation-Centric Evaluation Framework for 3D Medical Vision Foundation ModelsMaregu Assefa, Muhammad Muzammal Naseer, Divya Velayudhan, Gedamu Kumie and 1 moreParis Poster Session 5, Fri, Dec 11, 11:30 AM–1:30 PM, Paris Poster Hall · Published 2026– ReadersNo votes yet1/20 AI panelreviewers recommend itReaders and the AI panel: vote on this paper to see what they said.Worth readingNot for meOnly vote on papers you've read. Sign in with GitHub to vote.AI panel: 1 of 20 reviewers recommend itlenient 1/5medium 0/10strict 0/5
91%Must read?Must readVote to see the scoreNeurIPS 2026Khalifa University of Science, TKhalifaMohamed bin Zayed University of LLM evaluation & benchmarksKaliBench: A Fine-Grained Benchmark for Cybersecurity Tool Use on Kali Linux with Runtime-Free Verifiable RewardsKaliBench evaluates natural-language-to-CLI translation for 1,642 Kali Linux cybersecurity tools, finding open-weight models below 42% accuracy but training with verifiable rewards significantly improves smaller models.Pengfei Li, Naufal Suryanto, Sicheng Zhang, Muhammad Muzammal NaseerSydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026– ReadersNo votes yet17/20 AI panelreviewers recommend itReaders and the AI panel: vote on this paper to see what they said.Worth readingNot for meOnly vote on papers you've read. Sign in with GitHub to vote.AI panel: 17 of 20 reviewers recommend itlenient 5/5medium 8/10strict 4/5