
Sherpa: Teaching LLMs to Teach Adaptively
Sherpa uses multi-turn reinforcement learning to train LLM teachers that adapt instructions to diverse student archetypes, improving student performance by 20.5 points and pedagogy scores to 79.2%.
Published Oct 6, 2026 · ▲ 3 on Hugging Face · Code
Readers and the AI panel: vote on this paper to see what they said.
Only vote on papers you've read. Sign in with GitHub to vote.











































