Good Papers

Showing papers from IBM Research Show all papers

67%Highly rated
?Highly ratedVote to see the score

Black-Box Uncertainty Quantification for Large Language Models via Ensemble-of-Ensembles

Wang Ma, Debarun Bhattacharjya, Junkyu Lee, Nhan H Pham and 2 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 1/5
medium 1/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Contrastive Retrieval Heads for Improved Attention-Based Reranking

Linh Tran, Yulong Li, Radu Florian, Stacy Patterson and 2 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
67%Highly rated
?Highly ratedVote to see the score

TerraMesh-Masks: Open‑Vocabulary Segmentation for Earth Observation

Benedikt Blumenstiel, Hugues Devimeux, Johannes Jakubik, Konrad Schindler

Paris Poster Session 2, Wed, Dec 9, 5:00 PM–7:00 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 2/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Explanations over Graphs: An Agent Architecture for IT Enterprise Diagnostic Tasks

Saurabh Jha, Rohan R. Arora, Bhavya Bhavya, Noah Zheutlin and 5 more

Atlanta Poster Session 1, Wed, Dec 9, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

ProbMedTOD: A Bayesian Network Guided Task-Oriented Dialogue System for Patient History Taking

Vishal Vivek Saley, Bhavesh Gurnani, Dinesh Raghu, Mausam

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Understanding Model Reprogramming: A Reachability and Relabeling Perspective

Zesheng Ye, Pin-Yu Chen, Feng Liu

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

A Simple Class-Agnostic Approach to Enhance Fair Adversarial Training

Erh-Chung Chen, Pin-Yu Chen, I-Hsin Chung, Che-Rung Lee

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
91%Must read
?Must readVote to see the score

The Sparsity Whisperer

Difference-informed pruning preserves output differences via difference-aware weight scoring, improving LLM sparsity over activation and reconstruction baselines at minimal cost.

Linghao Kong, Inimai Subramanian, Micah Adler, Dan Alistarh and 2 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
17/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 17 of 20 reviewers recommend it
lenient 5/5
medium 10/10
strict 2/5
92%Must read
?Must readVote to see the score

General Agent Evaluation

A systematic comparison of general agent architectures finds backbone choice dominates performance while architecture shifts results up to 12pp, and open models suffer generality sinks.

Elron Bandel, Asaf Yehudai, Lilach Edelstein, Yehoshua Sagron and 11 more

Paris Poster Session 5, Fri, Dec 11, 11:30 AM–1:30 PM, Paris Poster Hall · Published 2026 · ▲ 14 on Hugging Face · Code ★ 76

– ReadersNo votes yet
19/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 19 of 20 reviewers recommend it
lenient 5/5
medium 10/10
strict 4/5
74%Highly rated
?Highly ratedVote to see the score

Federated Concept-Based Models: Interpretable models with distributed supervision

Federated concept-based models aggregate distributed concept annotations across institutions, adapt architectures to evolving supervision, and enable interpretable inference for locally unavailable concepts while preserving privacy.

Dario Fenoglio, Arianna Casanova Flores, Francesco De Santis, Gabriele Dominici and 5 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 9 of 20 reviewers recommend it
lenient 5/5
medium 4/10
strict 0/5
91%Must read
?Must readVote to see the score

Safety Reconstructed: Generative Modeling via Masked Diffusion Builds Strong Safety Guardrails

LLaDA-Guard uses masked diffusion to score responses under each safety label and classify by difference, improving calibration, reducing over-defense, and enabling token-level risk localization with 60.7% prompt rewriting success.

Gert Lek, Abele Mălan, Chaoyi Zhu, Pin-Yu Chen and 2 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
17/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 17 of 20 reviewers recommend it
lenient 5/5
medium 10/10
strict 2/5
88%Must read
?Must readVote to see the score

Learning Rate Matters: Vanilla LoRA May Suffice for LLM Fine-tuning

Vanilla LoRA matches variant performance within 1-2% when learning rates are tuned, and differing optimal rates stem from Hessian eigenvalue variations.

Yu-Ang Lee, Ching-Yun Ko, Pin-Yu Chen, Mi-Yen Yeh

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026 · ▲ 7 on Hugging Face · Code ★ 13

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 4/5
medium 9/10
strict 2/5
74%Highly rated
?Highly ratedVote to see the score

WAXAL: A Large-Scale Multilingual African Language Speech Corpus

WAXAL introduces an open 1,250-hour ASR and 235-hour TTS speech corpus for 24 African languages to advance inclusive speech technology.

MohamedElfatih MohamedKhair, Emmanuel Asiedu Brempong, Subhashini Venugopalan, Abdoulaye Diack and 8 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026 · ▲ 4 on Hugging Face

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 9 of 20 reviewers recommend it
lenient 5/5
medium 3/10
strict 1/5
86%Must read
?Must readVote to see the score

DiagnosticIQ: A Benchmark for LLM-Based Industrial Maintenance Action Recommendation from Symbolic Rules

DiagnosticIQ benchmarks LLM recommendation of industrial maintenance actions from symbolic rules across 6,690 questions, finding frontier models match human experts but break under structural perturbation due to calibration failures rather than capability gaps.

Devin Y De Silva, Dhaval Patel, Christodoulos Constantinides, Shuxin Lin and 7 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 6 on Hugging Face

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 2/5
70%Highly rated
?Highly ratedVote to see the score

Hyperparameter Transfer for Dense Associative Memories

Derives explicit hyperparameter transfer rules for Dense Associative Memories and validates them against large-scale training.

Roi Holtzman, Dmitry Krotov, Boris Hanin

Paris Poster Session 5, Fri, Dec 11, 11:30 AM–1:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
4/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 4 of 20 reviewers recommend it
lenient 0/5
medium 4/10
strict 0/5
78%Highly rated
?Highly ratedVote to see the score

Revisiting the Generic Transformer: Deconstructing a Strong Baseline for Time Series Foundation Models

A generic patch Transformer achieves state-of-the-art zero-shot time series forecasting via simple training, with scaling and data ablations isolating key performance drivers.

Yunshi Wen, Wesley M Gifford, Chandra Reddy, Lam Nguyen and 2 more

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 4/5
medium 5/10
strict 2/5
80%Must read
?Must readVote to see the score

HSCO-Bench: An Agent-Driven End-to-End Hardware-Software Co-design Benchmark for Systems-on-Chip

HSCO-Bench evaluates LLM agents on end-to-end hardware-software co-design for SoCs, finding only two frontier models generate valid prototypes with suboptimal resource use.

Pei-Huan Tsai, Kuan-Lin Chiu, William Baisi, Pin-Yu Chen and 1 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 4/5
medium 7/10
strict 1/5
83%Must read
?Must readVote to see the score

Steering Externalities: Benign Activation Steering Unintentionally Increases Jailbreak Risk for Large Language Models

Benign activation steering vectors inadvertently multiply jailbreak risks by eroding safety guardrails and raising attack success rates above 80%.

Chen Xiong, Zhiyuan HE, Pin-Yu Chen, Ching-Yun Ko and 1 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 1/5