Towards Real-World Environment-Aware Zero-Shot Text-to-Speech Synthesis via Disentangled Audio Infilling

1University of Science and Technology of China  ·  2National Institute of Informatics, Japan

Abstract

Recent zero-shot text-to-speech (TTS) systems achieve remarkable naturalness and speaker similarity, but typically assume high-quality speaker prompts and either strip away or entangle the acoustic environment with speaker characteristics, which limits their applicability in real-world scenarios. In this paper, we present DAIEN-TTS, an environment-aware zero-shot TTS framework that disentangles and jointly models speech, background noise, and reverberation, enabling independent control over timbre and acoustic environment through separate speaker and environment prompts. Built upon the flow-matching-based F5-TTS, DAIEN-TTS employs a speech-environment separation module to decompose environmental speech into speech, noise, and reverberation components, which are then injected into the diffusion transformer to enable environment-aware generation. The system is first trained on simulated data constructed by mixing clean speech with noise and room impulse responses, with a cross-speaker conditioning strategy to suppress speaker information leakage from the environment branch, and then fine-tuned on real-world speech data to bridge the simulated-to-real domain gap. At inference, a triple classifier-free guidance mechanism enables fine-grained control over speech, noise, and reverberation, along with an SNR adaptation strategy to align the synthesized speech with the environment prompt. Experiments on both simulated and real-world test sets show that DAIEN-TTS generates environmental personalized speech with high naturalness, strong speaker similarity, and faithful noise and reverberation reproduction, while offering controllability beyond prior environment-aware TTS systems.

Section I

Environment-Robust Generation

Synthesizing clean speech with a silence environment prompt, under clean or noisy-reverberant speaker prompts.

Clean Speaker Prompt

Text Speaker Prompt Ground Truth F5-TTS DAIEN-TTS† DAIEN-TTS
Then why should they be surprised when they see one?
Her old man was Doc Mitchell.
He didn't consider mending the hole — the stones could fall through any time they wanted.

Simulated Environmental Speaker Prompt

Text Speaker Prompt Ground Truth F5-TTS DAIEN-TTS† DAIEN-TTS
The silty bottom of the gulf also contributes to the high turbidity.
It was named for Doctor Joseph Warren, Revolutionary War patriot.
This new material had stronger alloys in the iron.

Section II

Environment-Aware Generation

Synthesizing environmental personalized speech that reproduces the acoustic environment specified by the environment prompt.

Simulated Environmental Prompts

Text Speaker Prompt Environment Prompt Ground Truth F5-TTS DAIEN-TTS† DAIEN-TTS
The remaining singles failed to hit the dance chart.
He later worked as a journalist, and then became a novelist and playwright.
Palmer switched political parties throughout his life, starting out a Democrat.
Then why should they be surprised when they see one?
He has also made three solo albums.
This new material had stronger alloys in the iron.

Real-World Environmental Prompts

Text Speaker Prompt Environment Prompt F5-TTS DAIEN-TTS† DAIEN-TTS DAIEN-TTS-FT
One way led to the left and the other to the right-straight up the mountain.
Thus walk I through my wonderland While all the evening is atune, Beneath the cypress trees that stand Like candles to the barren moon.
I say you do know what this means, and you must tell us.

Section III

TCFG Controllability

Demonstrating independent control over noise and reverberation by varying TCFG scales.

Varying αnoise  (fixed αspeech=1, αreverb=6)

Ground Truth 1 : 0 : 6 1 : 3 : 6 1 : 6 : 6 1 : 9 : 6

Varying αreverb  (fixed αspeech=1, αnoise=3)

Ground Truth 1 : 3 : 0 1 : 3 : 3 1 : 3 : 6 1 : 3 : 9