Good Papers

Showing papers from Capital One Show all papers

57%Worth a look
?Worth a lookVote to see the score

Argus: A Cross-Regime Benchmark for the Transferability of Uncertainty Quantification in Computer-Use Agents

DIVAKE KUMAR, Sina Tayebati, Devashri Naik, Amanda Rios and 4 more

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
88%Must read
?Must readVote to see the score

MAEB: Massive Audio Embedding Benchmark

MAEB benchmarks 30 audio tasks across 100+ languages, finding no single model dominates and acoustic and linguistic skills trade off.

Adnan E Assadi, Isaac Chung, Chenghao Xiao, Roman Solomatin and 14 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 25 on Hugging Face

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 4/5
medium 8/10
strict 3/5
80%Must read
?Must readVote to see the score

AVSD: Adaptive-View Self-Distillation by Balancing Consensus and Teacher-Specific Privileged Signals

AVSD separates cross-view consensus from privileged residuals in multi-view self-distillation to adaptively supervise reasoning models, improving math and code benchmarks over single-view methods and GRPO.

Duy Nguyen, Hanqi Xiao, Archiki Prasad, Zaid Khan and 6 more

Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 4/5
medium 7/10
strict 1/5
89%Must read
?Must readVote to see the score

CoT-Guard: Small Models for Strong Monitoring

CoT-Guard, a 4B-parameter chain-of-thought monitor, detects hidden code-generation objectives via SFT and RL, outperforming larger models including GPT-5.

Nirav Diwan, Han Wang, Berkcan Kapusuzoglu, Ramin Moradi and 5 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 3/5
86%Must read
?Must readVote to see the score

Improving Consistency in Retrieval Augmented Systems With Group Similarity Rewards

A framework decomposes RAG consistency into retrieval, generation, and end-to-end components, and PS-GRPO group similarity rewards train Con-RAG to boost consistency and accuracy across paraphrased queries.

Faisal Hamman, Chenyang Zhu, Anoop Kumar, Xujun Peng and 3 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 0/5
80%Must read
?Must readVote to see the score

SEAD: Competence-Aware On-Policy Distillation via Entropy-Guided Supervision

SEAD uses entropy-guided supervision at token, phase, and prompt levels to cut wasteful gradients and achieves +4.8 average accuracy over vanilla on-policy distillation.

Michael Lee, Zelei Cheng, Yu Wang, Renkun Ni and 3 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 3/5
medium 9/10
strict 0/5