Good Papers

Showing papers from International Business Machines Show all papers

90%Must read

A Matter of TASTE: Improving Coverage and Difficulty of Agent Benchmarks

TASTE reverses benchmark construction by evolving tool sequences to automatically generate harder, broader-coverage agent tasks that expose severe performance drops and saturation in existing benchmarks.

Tomer Keren, Nitay Calderon, Asaf Yehudai, Yotam Perlitz and 2 more

Paris Poster Session 6, Fri, Dec 11, 2:30 PM–4:30 PM, Paris Poster Hall · Published 2026 · ▲ 74 on Hugging Face · Code ★ 4

100% Readers1 of 1 upvoted
15/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 4/5
medium 9/10
strict 2/5
57%Worth a look
?Worth a lookVote to see the score

RAD-TFM: Robust and Domain-Adapted Tabular Foundation Models

Matthew Peroni, Franck Le, Vadim Sheinin

Atlanta Poster Session 4, Thu, Dec 10, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Explanations over Graphs: An Agent Architecture for IT Enterprise Diagnostic Tasks

Saurabh Jha, Rohan R. Arora, Bhavya Bhavya, Noah Zheutlin and 5 more

Atlanta Poster Session 1, Wed, Dec 9, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

A Simple Class-Agnostic Approach to Enhance Fair Adversarial Training

Erh-Chung Chen, Pin-Yu Chen, I-Hsin Chung, Che-Rung Lee

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
91%Must read
?Must readVote to see the score

SynBench: A Benchmark for Differentially Private Text Generation

SynBench benchmarks differentially private text generators across standardized datasets, revealing quality drops on out-of-distribution private data and invalidated privacy guarantees from pre-training contamination.

Yidan Sun, Viktor Schlegel, Srinivasan Nandakumar, Iqra Zahid and 8 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
17/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 17 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 4/5
88%Must read
?Must readVote to see the score

GIST: Gauge-Invariant Spectral Transformers for Scalable Graph Neural Operators

GIST proposes gauge-invariant spectral transformers that use efficient spectral embeddings to achieve linear complexity and provable discretization-invariance, setting state-of-the-art on large-scale mesh benchmarks.

Mattia Rigotti, Nicholas Thumiger, Thomas Frick

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 3/5
medium 10/10
strict 2/5
92%Must read
?Must readVote to see the score

General Agent Evaluation

A systematic comparison of general agent architectures finds backbone choice dominates performance while architecture shifts results up to 12pp, and open models suffer generality sinks.

Elron Bandel, Asaf Yehudai, Lilach Edelstein, Yehoshua Sagron and 11 more

Paris Poster Session 5, Fri, Dec 11, 11:30 AM–1:30 PM, Paris Poster Hall · Published 2026 · ▲ 14 on Hugging Face · Code ★ 76

– ReadersNo votes yet
19/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 19 of 20 reviewers recommend it
lenient 5/5
medium 10/10
strict 4/5
88%Must read
?Must readVote to see the score

Learning Rate Matters: Vanilla LoRA May Suffice for LLM Fine-tuning

Vanilla LoRA matches variant performance within 1-2% when learning rates are tuned, and differing optimal rates stem from Hessian eigenvalue variations.

Yu-Ang Lee, Ching-Yun Ko, Pin-Yu Chen, Mi-Yen Yeh

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026 · ▲ 7 on Hugging Face · Code ★ 13

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 4/5
medium 9/10
strict 2/5
86%Must read
?Must readVote to see the score

DiagnosticIQ: A Benchmark for LLM-Based Industrial Maintenance Action Recommendation from Symbolic Rules

DiagnosticIQ benchmarks LLM recommendation of industrial maintenance actions from symbolic rules across 6,690 questions, finding frontier models match human experts but break under structural perturbation due to calibration failures rather than capability gaps.

Devin Y De Silva, Dhaval Patel, Christodoulos Constantinides, Shuxin Lin and 7 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 6 on Hugging Face

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 2/5
72%Highly rated
?Highly ratedVote to see the score

Wavefunction Flows: Efficient Quantum Simulation of Continuous Flow Models

Flow models relate to Schrödinger dynamics via an unusual Hamiltonian, enabling efficient quantum Hamiltonian simulation for preparing coherent encodings of flow-modeled distributions and supporting quantum statistical algorithms.

David Layden, Ryan Sweke, Vojtech Havlicek, Anirban Chowdhury and 1 more

Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
8/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 8 of 20 reviewers recommend it
lenient 2/5
medium 3/10
strict 3/5
83%Must read
?Must readVote to see the score

Steering Externalities: Benign Activation Steering Unintentionally Increases Jailbreak Risk for Large Language Models

Benign activation steering vectors inadvertently multiply jailbreak risks by eroding safety guardrails and raising attack success rates above 80%.

Chen Xiong, Zhiyuan HE, Pin-Yu Chen, Ching-Yun Ko and 1 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 1/5