Good Papers

Showing Instruction tuning Show all papers

78%Highly rated

Latent-MOPD: Latent Multi-Teacher On-Policy Distillation

Latent-MOPD distills multi-teacher LLM specialists via hidden-state and prediction-level on-policy supervision, outperforming token-only and representation-only baselines across math, code, and logic benchmarks.

Zhengyu Fang, Seoyeon Hong, Jie Yang, Muyang Li and 3 more

Published Oct 1, 2026 · 0 citations · ▲ 63 on Hugging Face

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 3/5
medium 7/10
strict 1/5
78%Highly rated
?Highly ratedVote to see the score

Local Support Learning

Local Support Learning pairs weight adapters with GMM gating to keep updates local, resolving catastrophic forgetting in LLMs up to 7B parameters without prior data.

Assaf Ben-Kish, Akarsh Kumar, James Glass, Raja Giryes

Published Oct 1, 2026 · 0 citations · ▲ 25 on Hugging Face · Code ★ 13

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 4/5
medium 7/10
strict 0/5
80%Must read
?Must readVote to see the score

LLM2Jev: LLMs Are Already Jev-Style Decision Models -- When and How to Fine-Tune Them

LLM2Jev extracts calibrated Jev-style decisions from LLM token probabilities via training-free inference or tree-factorized fine-tuning, showing strong 4B models already match specialized decision models while fine-tuning mainly helps weaker backbones and specific tasks without degrading generation.

Yinheng Li, Justin Wagle

Published Oct 1, 2026 · 0 citations · ▲ 6 on Hugging Face

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 5/5
medium 6/10
strict 1/5
76%Highly rated
?Highly ratedVote to see the score

FedFit: Federated Fine-Tuning of LLMs via Vector-Bank Parameterization and Quantization

FedFit reduces federated LLM fine-tuning overhead via vector-bank adapter parameterization and quantization, resolving LoRA aggregation conflicts to achieve up to 100x compression with comparable perplexity.

Hang Zou, Chao Zhang, Yuzhi Yang, Yu Tian and 2 more

Published Oct 1, 2026 · 0 citations

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 10 of 20 reviewers recommend it
lenient 4/5
medium 6/10
strict 0/5
89%Must read
?Must readVote to see the score

Rethinking Token Reweighting for SFT: Suppress, Reverse, and Extrapolate Learned Features

Existing token-reweighting methods cannot reverse harmful SFT features; SCALE uses frozen SFT deltas with entropy-guided gates to suppress, reverse, or extrapolate them, improving math and code results.

Cunchun Li, Haonan He, Yifan Gao, Minglei Li and 3 more

Published Sep 27, 2026 · 0 citations · ▲ 11 on Hugging Face

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 4/5
medium 10/10
strict 2/5
78%Highly rated
?Highly ratedVote to see the score

ScopeIF: Improving Scope-Aware Precise Instruction-Following in Large Language Models via Graded Reward Modeling

ScopeIF improves LLM instruction-following via graded reward modeling and scope-aware constraints, enabling small models to match frontier performance.

Bosi Wen, Yilin Niu, Xiaoying Ning, Ying Zhang and 2 more

Published Sep 26, 2026 · 0 citations

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 4/5
medium 7/10
strict 0/5
80%Must read
?Must readVote to see the score

From Personal to Collective: On the Role of Local and Global Knowledge in LLM Personalization

LoGo augments individual user signals with evolving global and community-level behavioral patterns via adaptive mediation, improving LLM personalization and reducing overfitting.

Zehong Wang, Junlin Wu, Zhaoxuan Tan, Bolian Li and 3 more

Published 2026 · 0 citations

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 0/5
69%Highly rated
?Highly ratedVote to see the score

Instant Personalized Large Language Model Adaptation via Hypernetwork

A hypernetwork enables instant personalized large language model adaptation by generating user-specific parameters directly from user data.

Zhaoxuan Tan, Zixuan Zhang, Haoyang Wen, Zheng Li and 7 more

Published 2026 · 1 citation

– ReadersNo votes yet. 1 from authors or colleagues not counted
3/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 3 of 20 reviewers recommend it
lenient 2/5
medium 1/10
strict 0/5
72%Highly rated
?Highly ratedVote to see the score

Instruction Tuning for Large Language Models: A Survey

Instruction tuning surveys supervised fine-tuning of LLMs on instruction-output pairs to align next-word prediction with human intent, covering datasets, training, applications, and limitations.

Shengyu Zhang, Linfeng Dong, Xiaoya Li, Sen Zhang and 7 more

Published Nov 17, 2025 · 79 citations

– ReadersNo votes yet
8/21 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 8 of 21 reviewers recommend it
lenient 5/5
medium 2/11
strict 1/5
71%Highly rated
?Highly ratedVote to see the score

Instruction Tuning for Story Understanding and Generation with Weak Supervision

Weak to Strong Instruction Tuning improves story understanding and generation by training models on instructions of varying clarity, outperforming state-of-the-art baselines.

Yangshu Yuan, Heng Chen, Christian Ng

Published Jan 26, 2025 · 0 citations

– ReadersNo votes yet
6/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 6 of 20 reviewers recommend it
lenient 4/5
medium 2/10
strict 0/5
89%Must read
?Must readVote to see the score

Stronger Models are NOT Stronger Teachers for Instruction Tuning

Stronger models are not stronger teachers for instruction tuning due to teacher-student incompatibility; a compatibility-adjusted reward metric predicts effective generators.

Zhangchen Xu, Fengqing Jiang, Luyao Niu, Lin, Bill Yuchen and 1 more

Published Nov 11, 2024 · 0 citations · ▲ 39 on Hugging Face

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 2/5
78%Highly rated
?Highly ratedVote to see the score

Model Swarms: Collaborative Search to Adapt LLM Experts via Swarm Intelligence

Model Swarms uses swarm intelligence to collaboratively adapt LLM experts in weight space, improving over baselines by up to 21% with minimal data and no tuning.

Shangbin Feng, Zifeng Wang, Yike Wang, Sayna Ebrahimi and 8 more

Published Oct 15, 2024 · 2 citations

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 4/5
medium 6/10
strict 1/5
71%Highly rated
?Highly ratedVote to see the score

QA-LoRA: Quantization-Aware Low-Rank Adaptation of Large Language Models

QA-LoRA proposes quantization-aware low-rank adaptation that quantizes LLM weights during fine-tuning and merges adapters into quantized models without accuracy loss.

Yuhui Xu, Lingxi Xie, Xiaotao Gu, Xin Chen and 5 more

Published Sep 26, 2023 · 21 citations · ▲ 46 on Hugging Face · Code ★ 148

– ReadersNo votes yet
6/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 6 of 20 reviewers recommend it
lenient 5/5
medium 1/10
strict 0/5
72%Highly rated
?Highly ratedVote to see the score

LoraHub: Efficient Cross-Task Generalization via Dynamic LoRA Composition

LoraHub dynamically combines existing LoRA modules without extra parameters or gradients to generalize to unseen tasks with few examples, trading some accuracy for much lower inference token costs versus in-context learning.

Chengsong Huang, Qian Liu, Lin, Bill Yuchen, Tianyu Pang and 2 more

Published Jul 25, 2023 · 7 citations · ▲ 34 on Hugging Face · Code ★ 668

– ReadersNo votes yet
8/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 8 of 20 reviewers recommend it
lenient 5/5
medium 3/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

SDS-LoRA: Overcoming Anisotropic Gradient Scaling in Low-Rank Adaptation

JungHun Oh, Sungyong Baik, Kyoung Mu Lee

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

LEAN: Library-Based Adaptation for Asynchronous, Federated Fine-Tuning

Erdong Hu, Yuxin Tang, Zhimin Ding, Christopher Jermaine

Paris Poster Session 3, Thu, Dec 10, 12:30 PM–2:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

ToPA: Block-wise Toeplitz Adaptation for Expressive and Efficient Fine-Tuning

Sicong Li, Qianqian Xu, Zhiyong Yang, Zitai Wang and 3 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
67%Highly rated
?Highly ratedVote to see the score

Adaptive Fine-Tuning Scheduler for Multi-Tenant Edge LLM via Convergence-Aware Bandits

Yandi Li, Jianxiong Guo, Yupeng Li, Zhiqing Tang and 2 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 2/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

VoluCore: Spanning Teacher Representations with Volumetric Coresets for Data-Efficient LLM Distillation

Wang Xi, Yue Wang

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Mask-Conditioned Gradient Masking for Fine-Tuning Mixture-of-Experts Diffusion Language Models

Yiru Tang, Kun Zhou, Xin Zhao, Jing Sha and 2 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
Show 20 more papers