Good Papers

Showing Data curation Show all papers

83%Must read
?Must readVote to see the score

DataFlow: An LLM-Driven Framework for Unified Data Preparation and Workflow Automation in the Era of Data-Centric AI

DataFlow is an LLM-driven framework providing modular data preparation pipelines and automated operator synthesis that improves downstream LLM performance over human-curated and synthetic baselines.

Liang, Hao, Ma, Xiaochen, Liu, Zhou, Wong, Zhen Hao and 31 more

Published Dec 18, 2025 · 0 citations · ▲ 225 on Hugging Face · Code ★ 8,195

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 1/5
57%Worth a look
?Worth a lookVote to see the score

GradTrack: Detecting Noisy Labels via Temporal Trajectories of Class-wise Gradient Misalignment

Townim Faisal Chowdhury, Ta Duc Huy, Hieu Phan, Anton van den Hengel and 3 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Subdata Selection: A Unified Framework for Optimal Selection and Statistical Efficiency Assessment

Min Yang, Wei Zheng, John Stufken, Ming-Chung Chang and 2 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Guiding Data Allocation for Robust Subpopulation Generalization

Zhaoying Pan, Yipei Wang, Shenyu Lu, Xiaoqian Wang

Atlanta Poster Session 2, Wed, Dec 9, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Readiness-Aware Sample Selection for Noisy Labels with Class Imbalance

Chihyeon Choi, Sangho Lee, Jiho Hong, Youngdoo Son and 1 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

PatchScout: Thematic Web Data Collection via Information Foraging

Michael West, Eduard Dragut

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
88%Must read
?Must readVote to see the score

FineVision: Open Data Is All You Need

FineVision unifies 24 million vision-language samples via rigorous curation and decontamination, and models trained on it outperform existing open mixtures across broad evaluations.

Luis Wiedmann, Orr Zohar, Amir Mahla, Xiaohan Wang and 5 more

Paris Poster Session 6, Fri, Dec 11, 2:30 PM–4:30 PM, Paris Poster Hall · Published 2026 · ▲ 81 on Hugging Face

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 2/5
74%Highly rated
?Highly ratedVote to see the score

Leveraging Data Symmetries to Select an Optimal Subset of Training Data under Label Noise

Exploiting data symmetries improves k-NN accuracy for selecting low-noise training subsets, yielding near-optimal performance despite high-dimensional label noise.

Kumar Shubham, Pavan Karjol, Kiran M K, Prathosh AP

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 9 of 20 reviewers recommend it
lenient 4/5
medium 5/10
strict 0/5