45%Niche pick?Niche pickVote to see the scoreNeurIPS 2026U California - Los AngelesU California, Los AngelesYaleMax Planck Institute for SoftwarMulti-agent LLM systemsYoudunit: Single-Call Counterfactual Necessity in Multi-Agent LLM SystemsMarissa Li, Stephanie Gao, Kenny Guo, Xingjian Li and 2 moreAtlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026– ReadersNo votes yet0/20 AI panelreviewers recommend itReaders and the AI panel: vote on this paper to see what they said.Worth readingNot for meOnly vote on papers you've read. Sign in with GitHub to vote.AI panel: 0 of 20 reviewers recommend itlenient 0/5medium 0/10strict 0/5
69%Highly rated?Highly ratedVote to see the scoreNeurIPS 2026MPI-SWSMax Planck Institute for BehavioMax Planck Institute for SoftwarLLM evaluation & benchmarksCan an LLM Reason Like a Lawyer? Benchmarking the ability of LLMs to map the facts of a case to the elements of the applicable legal ruleShounak Paul, Seungeon Lee, Christoph Engel, Krishna GummadiParis Poster Session 1, Wed, Dec 9, 12:30 PM–2:30 PM, Paris Poster Hall · Published 2026– ReadersNo votes yet3/20 AI panelreviewers recommend itReaders and the AI panel: vote on this paper to see what they said.Worth readingNot for meOnly vote on papers you've read. Sign in with GitHub to vote.AI panel: 3 of 20 reviewers recommend itlenient 2/5medium 1/10strict 0/5
78%Highly rated?Highly ratedVote to see the scoreNeurIPS 2026MPI-SWSMax Planck Institute for SoftwarKAISTRL for LLMsGeoX: Mastering Geospatial Reasoning Through Self-Play and Verifiable RewardsGeoX acquires geospatial reasoning via self-play with executable programs and verifiable rewards, improving base VLMs up to 5.5 points without large-scale human-curated data.Kyeongjin Ahn, Seungeon Lee, Krishna Gummadi, Meeyoung ChaParis Poster Session 6, Fri, Dec 11, 2:30 PM–4:30 PM, Paris Poster Hall · Published 2026– ReadersNo votes yet11/20 AI panelreviewers recommend itReaders and the AI panel: vote on this paper to see what they said.Worth readingNot for meOnly vote on papers you've read. Sign in with GitHub to vote.AI panel: 11 of 20 reviewers recommend itlenient 5/5medium 6/10strict 0/5