Good Papers

Cephalonauts One: A deep fMRI dataset for decoding naturalistic speech in the human brain

Cephalonauts One provides 30 hours per subject of whole-brain fMRI during naturalistic speech, paired with audio, transcripts, and embeddings, plus a brain decoding benchmark showing continuous performance gains with more training data.

Antoine Collas, Louis Jalouzot, Géraud Ilinca, Corentin Caris, Romain Valabregue, Ahmed Hassayoune, David Gonçalves, Madeleine Hueber, Thaddée Delebarre, Savatovsky Julien, Clara Fonteneau, Charles Maussion, Bertrand Thirion, Alexis Thual

Published 2026Paris Poster Session 2 · Wed, Dec 9, 5:00 PM–7:00 PM local time · Paris Poster HallarXiv ↗OpenReview ↗

71%
OverallHighly rated
?
OverallHighly ratedVote to see the scoreThe exact score shows once you've voted, so every vote is your own call. The first half of each home page shelf shows its scores.
Readers
–

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel7/20reviewers recommend it
lenient 4/5
medium 2/10
strict 1/5
AI panel?Vote to see what the 20 AI reviewers said

Abstract

Cephalonauts One is a whole-brain 3 Tesla (3T) functional magnetic resonance imaging (fMRI) dataset recorded while subjects listened to audio podcasts. Three healthy subjects underwent multiple scanning sessions, each consisting of five 15-minute runs, while listening to podcasts in their native language. With 30 hours of fMRI data per subject, the current release is the deepest available fMRI dataset using naturalistic speech stimuli. The dataset pairs brain activity with the corresponding podcast audio, transcript annotations, and derived stimulus embeddings. Furthermore, we introduce a brain decoding benchmark formulated as audio segment retrieval: given fMRI activity from a held-out session, the decoder must identify the corresponding time-aligned podcast audio segment among candidate segments. We provide standardized splits, evaluation metrics, and baseline decoders for this task. Finally, a scaling analysis shows that decoding performance improves continuously with the amount of training data per subject.