Good Papers

Showing papers from Washington University in St. Louis Show all papers

89%Must read
?Must readVote to see the score

RRSI: Regularized Recursive Self-Improvement of Agent Harnesses

RRSI regularizes recursive agent harness self-improvement via annealed edit budgets, trajectory exploration, and critical selection to boost out-of-distribution performance and reduce token use. It improves up to 14.1 points in-distribution and 4.7 points out-of-distribution while cutting policy tok

Peng Xia, Rujun Han, Zifeng Wang, Yanfei Chen and 10 more

Published Sep 21, 2026 · 0 citations · ▲ 222 on Hugging Face · Code ★ 1,293

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 2/5
88%Must read
?Must readVote to see the score

EnvHarness: Awakening Static Worlds for Agent Learning

EnvHarness wraps static environments with programmable components to reshape agent behavior without altering underlying logic, improving benchmarks by up to 9.0 points while enabling continuous policy-environment co-evolution.

Chengsong Huang, Zifeng Wang, Rujun Han, Jun Yan and 13 more

Published Aug 20, 2026 · 0 citations · ▲ 175 on Hugging Face · Code ★ 619

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 1/5
70%Highly rated
?Highly ratedVote to see the score

Small RL Controller, Large Language Model: RL-Guided Adaptive Sampling for Test-Time Scaling

A lightweight RL controller formulates adaptive sampling as an MDP to balance LLM answer correctness, latency, and computation cost at test time.

Runpeng Dai, Tong Zheng, Rui Liu, Chengsong Huang and 1 more

Published Jun 2, 2026 · 0 citations · ▲ 13 on Hugging Face

– ReadersNo votes yet
5/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 5 of 20 reviewers recommend it
lenient 4/5
medium 1/10
strict 0/5
86%Must read
?Must readVote to see the score

Benchmark^2: Systematic Evaluation of LLM Benchmarks

Benchmark² evaluates LLM benchmarks via ranking consistency, discriminability, and capability alignment deviation, revealing quality variations and enabling smaller effective test sets.

Qi Qian, Chengsong Huang, Jingwen Xu, Changze Lv and 12 more

Published Jan 7, 2026 · 0 citations · ▲ 34 on Hugging Face

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 1/5
80%Must read
?Must readVote to see the score

From Personal to Collective: On the Role of Local and Global Knowledge in LLM Personalization

LoGo augments individual user signals with evolving global and community-level behavioral patterns via adaptive mediation, improving LLM personalization and reducing overfitting.

Zehong Wang, Junlin Wu, Zhaoxuan Tan, Bolian Li and 3 more

Published 2026 · 0 citations

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 0/5
74%Highly rated
?Highly ratedVote to see the score

Divide, Reweight, and Conquer: A Logit Arithmetic Approach for In-Context Learning

LARA improves in-context learning by dividing long demonstrations into shorter parallel groups and reweighting their logits via non-gradient optimization, boosting accuracy and memory efficiency over baselines on BBH and MMLU.

Chengsong Huang, Langlin Huang, Jiaxin Huang

Published Oct 14, 2024 · 1 citation · ▲ 1 on Hugging Face · Code ★ 8

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 9 of 20 reviewers recommend it
lenient 4/5
medium 5/10
strict 0/5
72%Highly rated
?Highly ratedVote to see the score

LoraHub: Efficient Cross-Task Generalization via Dynamic LoRA Composition

LoraHub dynamically combines existing LoRA modules without extra parameters or gradients to generalize to unseen tasks with few examples, trading some accuracy for much lower inference token costs versus in-context learning.

Chengsong Huang, Qian Liu, Lin, Bill Yuchen, Tianyu Pang and 2 more

Published Jul 25, 2023 · 7 citations · ▲ 34 on Hugging Face · Code ★ 668

– ReadersNo votes yet
8/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 8 of 20 reviewers recommend it
lenient 5/5
medium 3/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Coarse-to-Fine 3D MRI Reconstruction via Resolution-Agnostic Neural Operators

Jiayun (Peter) Wang, Ruibo Wang, Valentin Duruisseaux, Ram Daftari and 4 more

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Skill-Level Effects in Behavioral Cloning: When Low-Skill Data Improves Performance

Saumik Narayanan, Kassa Korley, Siddhartha Sen, Chien-Ju Ho

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 1/5
45%Niche pick
?Niche pickVote to see the score

SLVR: Structured Latent Visual Reasoning via Human-like Reasoning Flows

Albert Gao, Bing Xue, Andrea Zanette

Atlanta Poster Session 2, Wed, Dec 9, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
74%Highly rated
?Highly ratedVote to see the score

Early Prediction of Future Behavioral Strategy from Process Traces

A process-level latent variable model predicts future behavioral strategy from partial cross-task process traces via transferable person-level representations. In PowerWash Simulator it predicts zone planner versus hopper behavior in held-out levels.

Robert Kasumba, Dennis Barbour, Chien-Ju Ho

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 9 of 20 reviewers recommend it
lenient 5/5
medium 3/10
strict 1/5
72%Highly rated
?Highly ratedVote to see the score

G-Zero: Self-Play for Open-Ended Generation from Zero Data

G-Zero uses intrinsic predictive-shift rewards in a verifier-free co-evolutionary framework that enables continuous LLM self-improvement across open-ended unverifiable domains without external judges.

Chengsong Huang, Haolin Liu, Tong Zheng, Runpeng(Leo) Dai and 6 more

Atlanta Poster Session 2, Wed, Dec 9, 4:30 PM–7:30 PM, Hall C1 · Published 2026 · ▲ 17 on Hugging Face · Code ★ 30

– ReadersNo votes yet
8/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 8 of 20 reviewers recommend it
lenient 3/5
medium 5/10
strict 0/5
89%Must read
?Must readVote to see the score

Scaling Storm-Resolving Atmospheric AI Simulation to the Entire Planet

STRATA is an autoregressive AI emulator for global storm-resolving atmospheric dynamics that achieves 50× better energy efficiency and stable 24-hour rollouts on limited training data.

Zeyuan Hu, Noah Brenowitz, Akshay Subramaniam, Jaideep Pathak and 8 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 3/5
74%Highly rated
?Highly ratedVote to see the score

UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation

UniVL embeds visual and textual instructions into spatial masks for contextual image generation, cutting FID to 11 and inference costs by 52% without a text encoder.

Jiayun (Peter) Wang, Yu Wang, Weijie Gan, Zhenting Wang and 1 more

Atlanta Poster Session 4, Thu, Dec 10, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 9 of 20 reviewers recommend it
lenient 4/5
medium 5/10
strict 0/5
80%Must read
?Must readVote to see the score

ActWorld: From Explorable to Interactive World Model via Action-Aware Memory

ActWorld extends interactive world models to object interaction via a 100K dataset and hierarchical action-aware memory, improving fidelity over navigation-only baselines.

Zhexiao Xiong, Yizhi Song, Hao Kang, Qing Yan and 9 more

Atlanta Poster Session 2, Wed, Dec 9, 4:30 PM–7:30 PM, Hall C1 · Published 2026 · ▲ 8 on Hugging Face

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 4/5
medium 7/10
strict 1/5
71%Highly rated
?Highly ratedVote to see the score

Adapting Actively on the Fly: Relevance-Guided Online Meta-Learning with Latent Concepts for Geospatial Discovery

A geospatial discovery framework combines active learning, online meta-learning, and concept relevance to robustly find hidden targets like PFAS under sparse, changing conditions with limited sampling budgets.

Jowaria Khan, Anindya Sarkar, Yevgeniy Vorobeychik, Elizabeth Bondi-Kelly

Atlanta Poster Session 5, Fri, Dec 11, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
7/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 7 of 20 reviewers recommend it
lenient 5/5
medium 2/10
strict 0/5