Good Papers

Showing Datasets & benchmarks Show all papers

71%Highly rated
?Highly ratedVote to see the score

BanglaDial-Abuse: A Corpus-Grounded Dataset for Regional Dialect Identification in Abusive Bangla Text

BanglaDial-Abuse introduces 1,000 synthetic abusive Bangla sentences across four regional dialects for four-class dialect identification, achieving 0.37, 0.56 lexical Jaccard similarity with distinct lexical spaces.

Hasin Almas Sifat

Published Oct 1, 2026 · 0 citations

– ReadersNo votes yet
7/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 7 of 20 reviewers recommend it
lenient 4/5
medium 2/10
strict 1/5
57%Worth a look
?Worth a lookVote to see the score

PULSE: A Synchronized Five-Modality Dataset for Sensorimotor Coordination in Long-Horizon Daily Activities

Wolin Liang, Yang Gao, Peiyu Yan, YueXiang Hu and 5 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Dataset Collections: Challenges of Large-Scale Data Aggregation in 3D Medical Image Datasets

Yannick Kirchhoff, Saikat Roy, Elisa Stegmeier, Hamideh Haghiri and 29 more

Paris Poster Session 5, Fri, Dec 11, 11:30 AM–1:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

spora: A Unified Multimodal Dataset for Spatial Proteomics

Benedikt von Querfurth, Eeshaan Jain, Johann Wenckstern, Lukas Klein and 8 more

Paris Poster Session 6, Fri, Dec 11, 2:30 PM–4:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
67%Highly rated
?Highly ratedVote to see the score

SafeDrug: A Benchmark Dataset for Safety-Critical Pharmacological Reasoning in LLMs

Tengfei Ma, Yushan Yang, HOU Jiahao, Yujie Chen and 5 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 2/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

WirelessMathBench-XL: A Contamination-Audited Benchmark for Wireless Mathematical Reasoning

Xin Li, Mengbing Liu, Yiyang Zhu, WENHE ZHANG and 3 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

ECLIPSE: A Spacecraft Rendezvous Trajectories Dataset with Controlled In-Orbit Lighting Conditions

Nidhal Eddine Chenni, Arunkumar Rathinam, Abid Ali, Djamila Aouada

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

SynthHair: Leveraging MetaHumans for a High-Quality 4K Hair Matting Dataset

Markus Karmann, Shile Li, Philip Torr, Puneet Dokania and 4 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

The BAMBI Dataset: Multimodal Nadir UAV-Recordings of Forest Wildlife

Christoph Praschl, Hugo Markoff, Anna Maschek, Wolfram Jantsch and 7 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

MySign: A High-Fidelity Motion-Capture Dataset for 3D Sign Generation in Bahasa Isyarat Malaysia

Jiayu Shen, Kalin Stefanov, Lay-Ki Soon, Vee Yee CHONG and 2 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

GameVerse: A Minute-Scale Gameplay Dataset for Long-Horizon Interactive World Modeling

Kang He, Wenshuo Peng, Chuanhao Li, Zihui Gao and 6 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

MathCD: A Benchmark Dataset for Cognitive Diagnosis with Semantic Information

Xueyi Li, Youheng Bai, Tengteng Cheng, Mingliang Hou and 4 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

OTel: Open Telco AI Datasets, Benchmarks, and Models

Farbod Tavakkoli, Gregory Diamos, Kenneth Church, David Kanter and 14 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

TabBioMed: A Large-Scale Benchmark for Biomedical Tabular Learning

Pau Mateo Bernadó, Saivenkata Nagavyjayanthi Polapragada, Pol Arbiol Rakuljic, Laia M Pladevall and 7 more

Paris Poster Session 4, Thu, Dec 10, 5:30 PM–7:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

MEV: A Multi-Event Video Dataset for Long-Take Generation

Peiyuan Zhu, Shaoan Xie, Yifan Shen, Wen Tian and 3 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Tesserae and MoDiCo: A Billion-Fragment Dataset and Multi-Branch Architecture for File Fragment Classification

James Ghawaly

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Benchmark Shadows: How Data Regimes Shape Parameter Footprints and Generalization

Hongjian Zou, Yidan Wang, dingqi, Yixuan Liao and 1 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

LSC-Parlament: An Automatically Aligned Catalan Sign Language Dataset from Parliament Videos.

Carlos Escolano, Gerard Sant, Marc J Garcia, Joan G Cortés and 3 more

Paris Poster Session 5, Fri, Dec 11, 11:30 AM–1:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Autonomous Driving Research Requires a Community-Driven Data Paradigm

Jinsu Yoo, Zanming Huang, Katie Luo, Zheda Mai and 4 more

Atlanta Poster Session 5, Fri, Dec 11, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

RoMo Hands: A Large Scale Richly Organized Text to Hand Motion Dataset

Yizhak Ben-Shabat, Jiahao Zhang, Joseph Liu, Seonghyeon Moon and 4 more

Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 1/5
57%Worth a look
?Worth a lookVote to see the score

Ground False: Uncovering Errors in Formal Mathematics Benchmarks

Marcus Min, One An, Xujie Si, Osbert Bastani

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 1/5
67%Highly rated
?Highly ratedVote to see the score

GPT-Image-Edit-1M: An Auditable Million-Scale Dataset for Instruction-Guided Image Editing

Yuhan Wang, Siwei Yang, Bingchen Zhao, Letian Zhang and 7 more

Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 1/5
57%Worth a look
?Worth a lookVote to see the score

CounterStrike-1K: A Multi-Perspective Dataset of Professional Gameplay for World Modeling

Anirudhh Ramesh

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
76%Highly rated
?Highly ratedVote to see the score

The Road Ahead in Autonomous Driving: The KITScenes Multimodal Dataset

KITScenes Multimodal provides European autonomous driving data with high-fidelity synchronized sensors, complete topologically connected 3D HD maps, and four embodied AI benchmarks.

Richard Schwarzkopf, Fabian Immel, Alexander Blumberg, Jonas Merkert and 20 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 19 on Hugging Face · Code ★ 26

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 10 of 20 reviewers recommend it
lenient 5/5
medium 3/10
strict 2/5
80%Must read
?Must readVote to see the score

Croissant Baker: Metadata Generation for Discoverable, Governable, and Reusable ML Datasets

Croissant Baker generates validated Croissant metadata locally from dataset directories via modular handlers, achieving 97, 100% agreement with ground truth across 140+ datasets including MIMIC-IV.

Rafi Al Attrach, Rajna Fani, Sebastian Lobentanzer, Joan Giner-Miguelez and 16 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 5/5
medium 5/10
strict 2/5
83%Must read
?Must readVote to see the score

A Matched-Budget Audit Framework for Recaptioned Image-Text Supervision Distributions

A matched-budget audit framework profiles recaptioned image-text distributions via five axes and controllable basic units, showing released captions increase supported units by 3.39 to 6.36 and revealing a long-vs-dense frontier.

Giyeong Oh, Junghun Park, Yuhan Bae, Youngjae Yu

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 3/5
medium 9/10
strict 1/5
78%Highly rated
?Highly ratedVote to see the score

VAANI: Capturing the language landscape for an inclusive digital India

Project VAANI releases a multimodal dataset of 31,255 speech hours and 289K images spanning 105 Indic languages across 165 Indian districts to support inclusive speech technology.

Sujith Pulikodan, Abhayjeet Singh, Agneedh Basu, Nihar Desai and 16 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 4 on Hugging Face

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 5/5
medium 5/10
strict 1/5
88%Must read
?Must readVote to see the score

AmaraSpatial-10K: A Spatially and Semantically Aligned 3D Dataset for Spatial Computing and Embodied AI

AmaraSpatial-10K is a 10,000 synthetic 3D asset dataset optimized for deployment, achieving 3.4x CLIP recall over Objaverse and 99.1% physics stability.

Mohammadsadegh Salehi, Alexander J Perkins, Igor P Maurell, Ashkan Dabbagh and 1 more

Paris Poster Session 1, Wed, Dec 9, 12:30 PM–2:30 PM, Paris Poster Hall · Published 2026 · ▲ 3 on Hugging Face

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 1/5
88%Must read
?Must readVote to see the score

Doomed to Re-Annotate, Forever: The ImageNet Story

ReImageNet reveals ~12% of ImageNet-1k labels are wrong and reannotation boosts top-1 accuracy up to 1.2% for supervised models and 5, 6% for MLLMs.

Illia Volkov, Nikita Kisel, Tetiana Mishkina, Klara Janouskova and 1 more

Paris Poster Session 2, Wed, Dec 9, 5:00 PM–7:00 PM, Paris Poster Hall · Published 2026 · ▲ 3 on Hugging Face

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 4/5
medium 8/10
strict 3/5
88%Must read
?Must readVote to see the score

EEG Benchmarking Needs a Task Specification Layer: NeuroDoc for Rulebook-Guided, Executable Benchmark Construction

NeuroDoc introduces a rulebook-guided task specification layer that standardizes EEG benchmarks into 53 reviewed entries with 245 executable task definitions across four model backbones.

Chengxuan Qin, 致格 陈, Pengshu, Rui Yang and 8 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 2/5
91%Must read
?Must readVote to see the score

Self Driving Datasets: From 20 Million Papers to Nuanced Biomedical Knowledge at Scale

PubMed is autonomously converted into structured biomedical datasets larger, more nuanced, and more accurate than manual repositories via ontology tagging, hybrid retrieval, and a multi-agent extraction system.

Haydn Jones, Yimeng Zeng, Alden Rose, Yifei Li and 10 more

Atlanta Poster Session 4, Thu, Dec 10, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
18/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 18 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 4/5
71%Highly rated
?Highly ratedVote to see the score

L-FAME: Longitudinal Focused Attention Meditation EEG Dataset and Benchmark

L-FAME offers a longitudinal EEG dataset and benchmark of 74 participants across three meditation practices over six weeks, with baseline classification and cross-session adaptation results.

Angqi Li, Basit R Syed, Hamzeh Alzweri, Taosheng Liu and 2 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
7/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 7 of 20 reviewers recommend it
lenient 5/5
medium 1/10
strict 1/5
83%Must read
?Must readVote to see the score

LibriBrain100: One Hundred Hours of Broad and Deep MEG Data for Neural Speech Decoding at Scale

LibriBrain100 provides over 100 hours of MEG speech-decoding data, showing deep within-subject recordings and broad multi-subject data improve noninvasive word classification.

Francesco Mantegna, Dulhan Jayalath, Gereon Elvers, Tasha Kim and 10 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026 · ▲ 4 on Hugging Face · Code ★ 15

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 5/5
medium 6/10
strict 2/5
78%Highly rated
?Highly ratedVote to see the score

Metropolis-Scale Road Network Datasets for Fine-Grained Urban Traffic Modeling

New metropolis-scale road network datasets with real connectivity and 5-minute speed and volume data expose scalability limits in traffic forecasting and motivate a simple efficient baseline.

Fedor Velikonivtsev, Oleg Platonov, Ekaterina Alimaskina, Gleb Bazhenov and 1 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 5/5
medium 5/10
strict 1/5