Good Papers

SENSE: Semantic Neural Speech Synthesis from Brain Dynamics via Spatial Graph Encoding

SENSE uses graph-based EEG encoding and semantic conditioning to synthesize speech from brain dynamics, outperforming baselines on acoustic and semantic metrics with minimal training subjects.

Jisoo Park, Seonghak Lee, Hyojin Park, Junseok Kwon

Published 2026Sydney Poster Session 5 · Thu, Dec 10, 10:00 AM–1:00 PM local time · Hall 1-4arXiv ↗OpenReview ↗

88%
OverallMust read
?
OverallMust readVote to see the scoreThe exact score shows once you've voted, so every vote is your own call. The first half of each home page shelf shows its scores.
Readers
–

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel15/20reviewers recommend it
lenient 5/5
medium 8/10
strict 2/5
AI panel?Vote to see what the 20 AI reviewers said

Abstract

Reconstructing speech from non-invasive brain signals offers a promising pathway for restoring communication in individuals who are cognitively intact but unable to speak. Existing EEG-to-speech approaches formulate this task as acoustic reconstruction, optimizing waveform fidelity while ignoring whether the generated speech preserves high-level semantic content. In this work, we revisit this formulation and argue that EEG signals carry not only acoustic but also semantic information. We identify two key limitations of prior methods: (1) the neglect of spatial relationships between EEG electrodes, and (2) the failure to exploit the semantic structure of the N400 paradigm, where congruent and incongruent trials reflect distinct semantic processing. We propose SENSE(Semantic-EEG Neural Speech SynthEsis), which combines a graph-based EEG encoder over electrode geometry with EEG Semantic Conditioning (ESC), aligning EEG to a pretrained semantic space using only congruent trials. On the N400 dataset, SENSE consistently outperforms prior methods on both acoustic and semantic metrics, and model-internal channel attribution suggests distributed reliance on auditory, sensorimotor, and centro-parietal regions, consistent with known speech-perception neuroscience. In the unseen-subject setting, SENSE trained on only two subjects already surpasses the strongest baseline trained on all eighteen subjects in word error rate, and matches it on acoustic metrics with as few as eight subjects.