使用edge-tts与sounddevice播放时仅出短促爆音而非语音的问题求助
问题:edge-tts流式播放出现短促爆音而非正常语音
我正在编写一个脚本,计划通过edge-tts将文本转换为语音,再使用sounddevice进行流式播放,以便edge-tts生成完成后立即开始播放。但运行时仅听到短促(短于预期时长)的巨大爆音,而非正常语音,不清楚原因。以下是我的代码:
EDGE_SR = 24000 async def stream_speak(text, voice="de-DE-SeraphinaMultilingualNeural"): communicate = edge_tts.Communicate(text, voice=voice) device_info = sd.query_devices(kind='output') target_sr = int(device_info['default_samplerate']) print(f"Playing at {target_sr} Hz (resampling from {EDGE_SR} Hz)") with sd.OutputStream(24000, blocksize=1024, channels=1, dtype="int16") as stream: async for chunk in communicate.stream(): if chunk["type"] == "audio": audio = np.frombuffer(chunk["data"], dtype=np.int16) if target_sr != EDGE_SR: audio = resample_poly(audio, target_sr, EDGE_SR).astype(np.int16) stream.write(audio) def speak_sync(text): asyncio.run(stream_speak(text)) speak_sync('Hallo, ich kann dich gut hören')
请问有人知道问题出在哪吗?恳请提供帮助。
解决方案
导致爆音和播放异常的核心问题有两个,修正后即可正常播放:
采样率不匹配:你在
sd.OutputStream中硬编码了采样率为24000,但实际获取的设备默认采样率是target_sr。如果两者不一致,播放时音频的速度会被错误拉伸,直接导致爆音和时长异常。必须将OutputStream的采样率参数改为target_sr。异步上下文同步IO阻塞风险:
stream.write()是同步方法,在异步循环中直接调用可能导致阻塞,影响流式播放的连贯性。可以将写入操作包装在异步线程池中执行,避免阻塞事件循环。
修正后的代码如下:
import asyncio import numpy as np import sounddevice as sd from scipy.signal import resample_poly import edge_tts EDGE_SR = 24000 async def stream_speak(text, voice="de-DE-SeraphinaMultilingualNeural"): communicate = edge_tts.Communicate(text, voice=voice) device_info = sd.query_devices(kind='output') target_sr = int(device_info['default_samplerate']) print(f"Playing at {target_sr} Hz (resampling from {EDGE_SR} Hz)") # 使用target_sr初始化OutputStream,匹配设备采样率 with sd.OutputStream(samplerate=target_sr, blocksize=1024, channels=1, dtype="int16") as stream: async for chunk in communicate.stream(): if chunk["type"] == "audio": audio = np.frombuffer(chunk["data"], dtype=np.int16) if target_sr != EDGE_SR: # 重采样时确保计算正确,resample_poly的参数是(目标采样率, 原采样率) audio = resample_poly(audio, target_sr, EDGE_SR).astype(np.int16) # 使用线程池执行同步写入,避免阻塞异步循环 await asyncio.get_event_loop().run_in_executor(None, stream.write, audio) def speak_sync(text): asyncio.run(stream_speak(text)) speak_sync('Hallo, ich kann dich gut hören')
额外注意事项:
- 确保
scipy库已正确安装(resample_poly依赖它),如果不想依赖scipy,也可以使用其他轻量重采样库如librosa的resample方法。 - 部分设备可能对
blocksize参数敏感,如果仍有异常,可以尝试调整blocksize为2048或4096,找到最适合设备的数值。
内容的提问来源于stack exchange,提问作者Artem Melnyk
相关产品推荐
相关产品推荐

