Good Papers

Showing papers from Google, Tel Aviv University Show all papers

71%Highly rated
?Highly ratedVote to see the score

The FACTS Leaderboard: A Comprehensive Benchmark for Large Language Model Factuality

The FACTS Leaderboard benchmarks large language model factuality across multimodal, parametric, search, and grounding tasks via automated judges.

Aileen Cheng, Alon Jacovi, Amir Globerson, Ben Golan and 36 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 8 on Hugging Face

– ReadersNo votes yet
7/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 7 of 20 reviewers recommend it
lenient 4/5
medium 2/10
strict 1/5
88%Must read
?Must readVote to see the score

Controllable User Simulation

Controllable user simulation is formalized as causal inference, proving supervised fine-tuning injects look-ahead bias causing geometric variance explosion and controllability collapse, with proposed mitigations restoring consistency and robust generalization.

Guy Tennenholtz, Ofer Meshi, Amir Globerson, Uri Shalit and 2 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 1 on Hugging Face

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 4/5
medium 7/10
strict 4/5
71%Highly rated
?Highly ratedVote to see the score

Online Differentially Private Consistent Clustering

Differentially private online clustering transforms streams into private semi-coresets via a generic reduction, matching or improving approximation, space, and runtime while inheriting consistency from underlying non-private algorithms.

Edith Cohen, Vadym Doroshenko, Badih Ghazi, Pritish Kamath and 6 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
6/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 6 of 20 reviewers recommend it
lenient 2/5
medium 4/10
strict 0/5