
OPUS: Towards Efficient and Principled Data Selection in Large Language Model Pre-training in Every Iteration
OPUS defines optimizer-induced update-space data utility for dynamic LLM pre-training selection, outperforming full-scale baselines with minimal overhead.
Published Feb 5, 2026 · 0 citations · ▲ 354 on Hugging Face
Readers and the AI panel: vote on this paper to see what they said.
Only vote on papers you've read. Sign in with GitHub to vote.

