Good Papers

Showing papers from Salesforce Research Show all papers

45%Niche pick
?Niche pickVote to see the score

SkillOrchestra: Learning to Route Agents via Skill Transfer

Jiayu Wang, Yifei Ming, Zixuan Ke, Shafiq Joty and 2 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
67%Highly rated
?Highly ratedVote to see the score

Risks Create a Jagged Frontier of LLM Productivity Gains Across Computer Occupations

Deepika Chawla, Gagandeep Singh, Elham k buxton, Meicen Sun and 3 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 2/5
medium 0/10
strict 0/5
88%Must read
?Must readVote to see the score

Prospective Hindsight: Self-Calibrating Reinforcement Learning via Prediction–Reality Gaps

Prospective Hindsight uses prediction-reality gaps to weight gradients, improving reinforcement learning performance and self-calibration by targeting blind spots without added objectives.

Jiaxin Zhang, XIANGYU PENG, Qinglin Chen, Yu Li and 2 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 2/5
86%Must read
?Must readVote to see the score

BehaviorBench: Modeling Real-World User Decisions from Behavioral Traces

BehaviorBench evaluates personalized decision modeling using real-world wallet traces across belief and trade prediction tasks, showing personalization improves beliefs more than trades and reveals model failure modes.

Liangwei Yang, Jielin Qiu, Zixiang Chen, Ming Zhu and 8 more

Atlanta Poster Session 5, Fri, Dec 11, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 1/5
83%Must read
?Must readVote to see the score

Procedural Memory Distillation: Online Reflection for Self-Improving Language Models

Procedural Memory Distillation extracts cross-episode strategy patterns into reusable procedural memory, co-evolving with the policy to improve reasoning benchmarks by up to 13.6% over SDPO.

Ye Liu, Srijan Bansal, Bo Pang, Yang Li and 5 more

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 4/5
medium 8/10
strict 1/5
78%Highly rated
?Highly ratedVote to see the score

Learning from Language Feedback via Variational Policy Distillation

Variational Policy Distillation co-evolves an adaptive teacher and student via variational EM to extract dense token-level guidance from language feedback, outperforming RLVR and self-distillation baselines on reasoning and code tasks.

Yang Li, Erik Nijkamp, Semih Yavuz, Shafiq Joty

Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026 · ▲ 12 on Hugging Face

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 5/5
medium 6/10
strict 0/5
86%Must read
?Must readVote to see the score

TerraVis: Towards Evaluation of World-Grounded Visual Consistency in Text-to-Image Generation via MLLM Workflows

TerraVis evaluates world-grounded visual consistency in generated images via MLLM workflows, correlating best with human judgments while revealing substantial failures in top models.

Shuai Fu, Jing Gu, Jian Zhou, Zicheng Duan and 2 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 1 on Hugging Face

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 4/5
medium 9/10
strict 1/5
86%Must read
?Must readVote to see the score

The Illusion of Multi-Agent Advantage

Automatic multi-agent systems consistently underperform single-agent chain-of-thought self-consistency despite up to 10x cost, revealing automated architectures suffer from bloat and misaligned complexity rather than true multi-agent benefits.

Prathyusha Jwalapuram, Hehai Lin, Chuyuan Li, Fangkai Jiao and 6 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026 · ▲ 1 on Hugging Face

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 4/5
medium 9/10
strict 1/5
80%Must read
?Must readVote to see the score

Memory Retrieval for Changing Preferences

A Bayes-factor utility framework selects memory turns by evidence of latent preference changes and regulates access accordingly. It outperforms embedding retrieval on preference-intensive long-context dialogue tasks.

Yuehan Qin, Li Li, Linxin Song, Jiate Li and 3 more

Atlanta Poster Session 5, Fri, Dec 11, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 5/5
medium 6/10
strict 1/5
78%Highly rated
?Highly ratedVote to see the score

Reward Modeling for Multi-Agent Orchestration

OrchRM uses self-supervised orchestration-level reward modeling to train multi-agent orchestrators, cutting token usage by 10x and boosting accuracy up to 8%.

King Yeung Tsang, Zihao Zhao, Vishal Venkataramani, Haizhou Shi and 4 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 5/5
medium 6/10
strict 0/5