Good Papers

Showing papers from Amazon Web Services Show all papers

45%Niche pick
?Niche pickVote to see the score

SPLICE: Structured Prompt Local Iterative Combinatorial Evolution

Dr. Anish Acharya, Phillip Studans, Amit Dhanda, Ninad V Rao and 3 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
69%Highly rated
?Highly ratedVote to see the score

RobustGenBench: A Benchmark for Robust Generalization to Adversarial and Common Perturbations, with Applications to Vision and Vision-Enabled Large Language Models

Maxime Heuillet, JONAS NGNAWE, Yann Pequignot, Rishika Bhagwatkar and 5 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
3/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 3 of 20 reviewers recommend it
lenient 2/5
medium 1/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Verify0: Can AI Agents Build Formally Verified Software Repositories?

Zhe Ye, Hantao Lou, Yuechun Sun, Peiyang Song and 7 more

Atlanta Poster Session 4, Thu, Dec 10, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 0/5
medium 1/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

nnTrace: Detecting and Localizing Silent Bugs in Distributed Training

Haitian Jiang, Shaowei Zhu, Zhen Zhang, Zhenyu Song and 4 more

Atlanta Poster Session 1, Wed, Dec 9, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
83%Must read
?Must readVote to see the score

Lean Refactor: Multi-Objective Controllable Proof Optimization via Agentic Strategy Search

Lean Refactor uses retrieval-augmented agentic strategy search to multi-objectively refactor Lean proofs, achieving over 70% token compression and up to 60% faster compilation with stronger version transfer.

Jialin Lu, Soonho Kong, Rodrigo Stehling, Kaiyu Yang and 3 more

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026 · ▲ 2 on Hugging Face

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 1/5
80%Must read
?Must readVote to see the score

CORVUS: Context Optimization and Reduction Via Underlying Synchronization for LLM Coding Agents

CORVUS decouples file reads from observations via synchronized registries, cutting input tokens by 9-50% and reasoning cycles by up to 37% while preserving pass rates.

Mingwei Zheng, David OBrien, Siwei Cui, Pardis Pashakhanloo and 3 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 4/5
medium 7/10
strict 1/5
89%Must read
?Must readVote to see the score

Known By Their Actions: Fingerprinting LLM Browser Agents via UI Traces

UI traces from LLM web agents identify underlying models with 96% F1 via passive JavaScript tracking, though randomized delays only partially mitigate fingerprinting.

William Gitta Lugoloobi, Samuele Marro, Jabez Magomere, Joss Wright and 1 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026 · Code ★ 7

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 3/5
70%Highly rated
?Highly ratedVote to see the score

s2n-bignum-bench: A practical benchmark for evaluating low-level code reasoning of LLMs

s2n-bignum-bench evaluates LLM theorem proving on verified industrial cryptographic assembly using HOL Light proof synthesis. It provides a challenging, practically relevant benchmark beyond competition mathematics.

Balaji Rao, Soonho Kong, Juneyoung Lee, Carlo Lipizzi

Atlanta Poster Session 1, Wed, Dec 9, 10:00 AM–1:00 PM, Hall C1 · Published 2026 · ▲ 3 on Hugging Face · Code ★ 5

– ReadersNo votes yet
5/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 5 of 20 reviewers recommend it
lenient 4/5
medium 1/10
strict 0/5