Good Papers

Showing papers from ETH Zurich and INSAIT, Sofia University Show all papers

72%Highly rated
?Highly ratedVote to see the score

Not All Proofs Are Equal: Evaluating LLM Proof Quality Beyond Correctness

ProofRank benchmarks LLM proof quality via conciseness, ease, simplicity, diversity, and adaptivity, revealing quality differences and trade-offs with correctness.

Ivo Petrov, Jasper Dekoninck, Dimitar I. Dimitrov, Martin Vechev

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
8/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 8 of 20 reviewers recommend it
lenient 4/5
medium 4/10
strict 0/5
86%Must read
?Must readVote to see the score

Making Open-Source Text LLM Watermarks Durable Against Merging

Merge-Adversarial Training embeds durable text watermarks into open-source LLM weights via adversarial distillation, maintaining high detection rates after model merging while preserving capabilities.

Luisa Scharff, Thibaud Gloaguen, Robin Staab, Martin Vechev

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 2/5
83%Must read
?Must readVote to see the score

Every Bit, Everywhere, All At Once: A Binomial Multibit LLM Watermark

A binomial multibit LLM watermark directly encodes every payload bit at each token via a stateful encoder, outperforming baselines on large payloads with high robustness and proposing per-bit confidence scoring.

Thibaud Gloaguen, Robin Staab, Mark Vero, Martin Vechev

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 0/5