如何在Python中用PyAudio播放Coqui TTS生成的NumPy数组音频
解决Coqui TTS生成的NumPy数组直接用PyAudio播放的问题
你遇到的问题核心是Coqui TTS输出的numpy数组格式和PyAudio的播放要求不匹配,再加上PyAudio只认字节数据而非数组。下面是具体解决步骤和修正后的代码:
问题根源
- Coqui TTS默认输出float32类型的numpy数组,取值范围在[-1, 1]之间,但你配置的PyAudio格式
paInt16需要的是16位整数(范围[-32768, 32767])。 - 手动设置的
CHANNELS=2和RATE=16000不一定和所选TTS模型的实际输出参数匹配,必须严格对齐。 - PyAudio的
stream.write()方法只接受字节对象,不能直接传入numpy数组。
解决步骤
- 获取模型真实音频参数:通过TTS实例的内置属性获取准确的采样率和通道数,不要硬编码。
- 转换数组格式:把float32数组缩放到int16的取值范围,再转成字节数据。
- 对齐PyAudio配置:用模型的真实参数初始化PyAudio流。
修正后的代码
from TTS.api import TTS import pyaudio import numpy as np # 初始化TTS模型 model_name = TTS.list_models()[0] tts = TTS(model_name) # 设置文本、发音人和语言 text = "King Charles III. King of the United Kingdom and 14 other Commonwealth realms. Prince Charles acceded to the throne on September 8 2022 upon the death of his mother, Queen Elizabeth II. He was the longest-serving heir apparent in British history and was the oldest person to assume the throne, doing so at the age of 73." speaker = tts.speakers[0] language = tts.languages[0] # 生成音频numpy数组 audio_np = tts.tts(text, speaker=speaker, language=language) audio_np = np.array(audio_np) # 确保输出为numpy数组 # 获取模型的真实音频参数 sample_rate = tts.synthesizer.output_sample_rate channels = tts.synthesizer.config.audio.num_channels # 将float32数组转换为PyAudio支持的int16字节数据 # 先缩放范围:[-1,1] → [-32768, 32767] audio_int16 = (audio_np * 32767).astype(np.int16) # 转换为字节对象 audio_bytes = audio_int16.tobytes() # 初始化PyAudio并播放 p = pyaudio.PyAudio() stream = p.open( format=pyaudio.paInt16, channels=channels, rate=sample_rate, output=True ) # 播放音频 stream.write(audio_bytes) # 清理资源 stream.stop_stream() stream.close() p.terminate()
关键改动说明
- 用
tts.synthesizer的内置属性获取真实的采样率和通道数,避免硬编码导致的参数不匹配。 - 添加了格式转换流程:先把float32音频数据缩放至int16的合法范围,再转成PyAudio能识别的字节对象。
- 确保
audio_np是标准numpy数组(部分场景下TTS可能返回列表,需手动转换)。
如果仍有问题,可以先打印audio_np.dtype、sample_rate、channels的值,确认参数是否正确。
内容的提问来源于stack exchange,提问作者rupertbj
相关产品推荐
相关产品推荐

