Good Papers

Showing papers from Alibaba Group (China) Show all papers

89%Must read
?Must readVote to see the score

OPUS: Towards Efficient and Principled Data Selection in Large Language Model Pre-training in Every Iteration

OPUS defines optimizer-induced update-space data utility for dynamic LLM pre-training selection, outperforming full-scale baselines with minimal overhead.

Shaobo Wang, Xuan Ouyang, Tianyi Xu, Yuzheng Hu and 8 more

Published Feb 5, 2026 · 0 citations · ▲ 354 on Hugging Face

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 2/5
67%Highly rated
?Highly ratedVote to see the score

A Data-Efficient Path to Multilingual LLMs: Language Expansion via Post-training PARAM𝛥 Integration into Upcycled MoE

The method expands multilingual LLMs via post-training PARAMΔ integration into upcycled MoE for data-efficient language acquisition.

Hao Zhou, Tianhao Li, Zhijun Wang, Shuaijie She and 5 more

Published 2026 · 0 citations

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 1/5
medium 1/10
strict 0/5
72%Highly rated
?Highly ratedVote to see the score

Instruction Tuning for Large Language Models: A Survey

Instruction tuning surveys supervised fine-tuning of LLMs on instruction-output pairs to align next-word prediction with human intent, covering datasets, training, applications, and limitations.

Shengyu Zhang, Linfeng Dong, Xiaoya Li, Sen Zhang and 7 more

Published Nov 17, 2025 · 79 citations

– ReadersNo votes yet
8/21 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 8 of 21 reviewers recommend it
lenient 5/5
medium 2/11
strict 1/5
71%Highly rated
?Highly ratedVote to see the score

OpenHands: An Open Platform for AI Software Developers as Generalist Agents

OpenHands is an open MIT-licensed platform for building AI software developers that evaluate agents on SWE-BENCH and WebArena benchmarks.

Xingyao Wang, Boxuan Li, Yufan Song, Frank F. Xu and 20 more

Published Jul 23, 2024 · 15 citations · ▲ 90 on Hugging Face · Code ★ 90,160

– ReadersNo votes yet
7/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 7 of 20 reviewers recommend it
lenient 5/5
medium 2/10
strict 0/5
76%Highly rated
?Highly ratedVote to see the score

ChatIE: Zero-Shot Information Extraction via Chatting with ChatGPT

ChatIE reframes zero-shot information extraction as multi-turn ChatGPT dialogue, surpassing some fully supervised models on several benchmark datasets.

Wei, Xiang, Xingyu Cui, Ning Cheng, Xiaobin Wang and 8 more

Published Feb 20, 2023 · 147 citations

– ReadersNo votes yet
10/21 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 10 of 21 reviewers recommend it
lenient 5/5
medium 5/11
strict 0/5