45%Niche pick?Niche pickVote to see the scoreNeurIPS 2026Scale AIU Minnesota, MinneapolisLLM agents & planningWayfinder: Adaptive Resource Routing from Agent CitationsMiguel Romero Calvo, George KarypisSydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026– ReadersNo votes yet0/20 AI panelreviewers recommend itReaders and the AI panel: vote on this paper to see what they said.Worth readingNot for meOnly vote on papers you've read. Sign in with GitHub to vote.AI panel: 0 of 20 reviewers recommend itlenient 0/5medium 0/10strict 0/5
45%Niche pick?Niche pickVote to see the scoreNeurIPS 2026North Carolina StatePennsylvania StateScale AISnapchatGraph neural networksIncAgg: Efficient Memory-Enhanced Graph Learning via Incremental AggregationXingyue Shi, Zhichao Hou, Jiahao Zhang, Suhang Wang and 3 moreAtlanta Poster Session 1, Wed, Dec 9, 10:00 AM–1:00 PM, Hall C1 · Published 2026– ReadersNo votes yet0/20 AI panelreviewers recommend itReaders and the AI panel: vote on this paper to see what they said.Worth readingNot for meOnly vote on papers you've read. Sign in with GitHub to vote.AI panel: 0 of 20 reviewers recommend itlenient 0/5medium 0/10strict 0/5
67%Highly rated?Highly ratedVote to see the scoreNeurIPS 2026U Pennsylvania, University of PeCenter for AI SafetyScale AIU Pennsylvania Aethra LabsLLM evaluation & benchmarksBeyond Truthfulness: Evaluating Honesty in Large Language ModelsRichard Ren, Arunim Agarwal, Mantas Mazeika, Cristina Menghini and 11 moreAtlanta Poster Session 1, Wed, Dec 9, 10:00 AM–1:00 PM, Hall C1 · Published 2026– ReadersNo votes yet2/20 AI panelreviewers recommend itReaders and the AI panel: vote on this paper to see what they said.Worth readingNot for meOnly vote on papers you've read. Sign in with GitHub to vote.AI panel: 2 of 20 reviewers recommend itlenient 1/5medium 1/10strict 0/5
86%Must read?Must readVote to see the scoreNeurIPS 2026U PennsylvaniaScale AIDyna RoboticsU Nevada, Las VegasU ArizonaAgent benchmarks & environmentsSWE Atlas: Benchmarking Coding Agents Beyond Issue ResolutionSWE Atlas benchmarks coding agents on codebase Q&A, test writing, and refactoring, finding frontier models lead but all struggle with edge cases and engineering quality.Mohit Raghavendra, Soham Dan, Miguel Romero Calvo, Yannis He and 11 moreSydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026– ReadersNo votes yet14/20 AI panelreviewers recommend itReaders and the AI panel: vote on this paper to see what they said.Worth readingNot for meOnly vote on papers you've read. Sign in with GitHub to vote.AI panel: 14 of 20 reviewers recommend itlenient 5/5medium 8/10strict 1/5
92%Must read?Must readVote to see the scoreNeurIPS 2026NorthwesternScaleAIScale AIReflection AIU California, BerkeleyAgent benchmarks & environmentsMCP-Atlas: A Large-Scale Benchmark for Tool-Use Competency with Real MCP ServersMCP-Atlas benchmarks LLM tool-use on 1,000 real-server tasks, finding frontier models reach 82.2% pass rates but 63.3% of failures are cognitive.Chaithanya Bandi, Razvan Dumitru, Ben Hertzberg, Divyansh Agarwal and 15 moreAtlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026 · ▲ 1 on Hugging Face– ReadersNo votes yet19/20 AI panelreviewers recommend itReaders and the AI panel: vote on this paper to see what they said.Worth readingNot for meOnly vote on papers you've read. Sign in with GitHub to vote.AI panel: 19 of 20 reviewers recommend itlenient 5/5medium 9/10strict 5/5