Good Papers

Showing Speech synthesis Show all papers

78%Highly rated
?Highly ratedVote to see the score

DEFINE: Exemplar-Guided Accent Control for Zero-Shot TTS

DEFINE decouples speaker identity and accent in zero-shot TTS via separate audio exemplars and a single guidance weight, generalizing accent control beyond training accents with high speaker similarity.

Ambuj Mehrish, Abhinaba Roy, Alex Ivanov, T. Ahmed and 1 more

Published Sep 26, 2026 · 0 citations · ▲ 32 on Hugging Face · Code ★ 2

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 4/5
medium 6/10
strict 1/5
88%Must read
?Must readVote to see the score

Accent Analogy Guidance: More Speaker Similarity at Equal Accent in Cross-Lingual Voice Cloning

Accent analogy guidance subtracts estimated accent directions for cross-lingual voice cloning, raising speaker similarity above identity-accent trade-off curves across several open TTS models.

Yoomee Cho, Jisun Lee

Published Sep 24, 2026 · 0 citations · ▲ 4 on Hugging Face · Code

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 2/5
57%Worth a look
?Worth a lookVote to see the score

MFlowAudio: Efficient Text-to-Audio Synthesis via Mamba-based Stateful Flow Matching

Hao Dai, Panyu Chen, Jagmohan Chauhan

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
67%Highly rated
?Highly ratedVote to see the score

RVCBench: Benchmarking Robustness of Voice Cloning Across Modern Audio Generation Models

Ruinan Jin, Xinting Liao, Hanlin Yu, Deval Pandya and 1 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 1/5
medium 1/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

StyleStream 2.0: Fast and Controllable Streaming Voice Style Conversion

Yisi Liu, Nicholas Lee, Gopala Anumanchipalli

Atlanta Poster Session 4, Thu, Dec 10, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Direct Conditioning of Audio Diffusion Transformers on fMRI Reveals Cortical Contributions to Sound Reconstruction

Matteo Ciferri, Tonio Weidler, Matteo Ferrante, Nicola Toschi and 1 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
67%Highly rated
?Highly ratedVote to see the score

Efficient Brain-to-Speech Decoding with Fixed-Delay Spiking Neural Networks

Aleksandra Wisniewska, Seo-Hyun Lee, Seong-Whan Lee

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 2/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Unbounded Streaming Text-To-Speech with Prefixed Sliding Window Attention

Théodor Lemerle, Diego Torres Guarin, Téo Guichoux, Nicolas Obin and 1 more

Paris Poster Session 2, Wed, Dec 9, 5:00 PM–7:00 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
80%Must read
?Must readVote to see the score

AudioCALM: Continuous Autoregressive Language Modeling for Universal Audio Generation

AudioCALM unifies speech, sound, and music generation via continuous autoregressive flow matching and asymmetric mixture-of-experts, matching specialized state-of-the-art performance.

Huadai Liu, Kaicheng Luo, Wen Wang, Qian Chen and 3 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 3/5
medium 8/10
strict 1/5
91%Must read
?Must readVote to see the score

DUET: Unified Dual-Space Emotion Control for Diffusion and Flow-Matching Driven Text-to-Speech

DUET enables plug-and-play emotion control for pretrained diffusion and flow-matching TTS by steering hidden states and guiding mel-spectra via a differentiable vocoder, surpassing supervised emotional baselines.

Xu Zhang, Longbing Cao, zhangkai wu

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
18/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 18 of 20 reviewers recommend it
lenient 5/5
medium 10/10
strict 3/5
78%Highly rated
?Highly ratedVote to see the score

Brain2voice 2.0: High-performance voice synthesis brain-computer interface

Brain2voice 2.0 synthesizes highly intelligible real-time voice from brain signals using a multimodal Transformer, cutting word error rates to 5.24%.

Maitreyee Wairagkar, Aparna Srinivasan, Nicholas S Card, Tyler Singer-Clark and 6 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 5/5
medium 4/10
strict 2/5
88%Must read
?Must readVote to see the score

SENSE: Semantic Neural Speech Synthesis from Brain Dynamics via Spatial Graph Encoding

SENSE uses graph-based EEG encoding and semantic conditioning to synthesize speech from brain dynamics, outperforming baselines on acoustic and semantic metrics with minimal training subjects.

Jisoo Park, Seonghak Lee, Hyojin Park, Junseok Kwon

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 2/5
74%Highly rated
?Highly ratedVote to see the score

Generating the Unheard: Phylogeny-Guided Latent Generation for Ancestral Sound Reconstruction

This framework generates ancestral bird vocalizations by inferring decodable VAE latents guided by phylogenetic traits, achieving genuine generation and naturalistic audio quality.

Tianyi Xu, Shrinaath Narasimhan, Evan Gorstein, Santiago Perea and 2 more

Atlanta Poster Session 5, Fri, Dec 11, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 9 of 20 reviewers recommend it
lenient 4/5
medium 5/10
strict 0/5
76%Highly rated
?Highly ratedVote to see the score

Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts

PlanAudio uses an autoregressive LLM with semantic latent chain-of-thought to synthesize unified speech and sound audio directly from free-form text, outperforming pipeline and unified baselines.

Yuyue Wang, Xihua Wang, Xin Cheng, Yijing Chen and 1 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 10 of 20 reviewers recommend it
lenient 4/5
medium 6/10
strict 0/5
78%Highly rated
?Highly ratedVote to see the score

Reconstructing the Vocal Tract with Differentiable Acoustic Simulation

A differentiable GPU acoustic simulator reconstructs vocal tract geometry from speech via gradient descent, enabling cross-lingual autoencoding and unpaired MRI reconstruction.

Eric Chen, Jin Woo Lee, Vincent Sitzmann

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 5/5
medium 4/10
strict 2/5
78%Highly rated
?Highly ratedVote to see the score

Kinetic-Optimal Scheduling with Moment Correction for Metric-Induced Discrete Flow Matching in Zero-Shot Text-to-Speech

A kinetic-optimal scheduler and moment correction improve metric-induced discrete flow matching, yielding GibbsTTS with best objective naturalness and strong speaker similarity in zero-shot text-to-speech.

Dong Yang, YIYI CAI, Haoyu Zhang, Yuki Saito and 1 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 3/5
medium 7/10
strict 1/5