You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用edge-tts与sounddevice播放时仅出短促爆音而非语音的问题求助

问题:edge-tts流式播放出现短促爆音而非正常语音

我正在编写一个脚本,计划通过edge-tts将文本转换为语音,再使用sounddevice进行流式播放,以便edge-tts生成完成后立即开始播放。但运行时仅听到短促(短于预期时长)的巨大爆音,而非正常语音,不清楚原因。以下是我的代码:

EDGE_SR = 24000

async def stream_speak(text, voice="de-DE-SeraphinaMultilingualNeural"):
    communicate = edge_tts.Communicate(text, voice=voice)

    device_info = sd.query_devices(kind='output')
    target_sr = int(device_info['default_samplerate'])
    print(f"Playing at {target_sr} Hz (resampling from {EDGE_SR} Hz)")

    with sd.OutputStream(24000, blocksize=1024, channels=1, dtype="int16") as stream:
        async for chunk in communicate.stream():
            if chunk["type"] == "audio":
                audio = np.frombuffer(chunk["data"], dtype=np.int16)
                if target_sr != EDGE_SR:
                    audio = resample_poly(audio, target_sr, EDGE_SR).astype(np.int16)
                stream.write(audio)

def speak_sync(text):
    asyncio.run(stream_speak(text))

speak_sync('Hallo, ich kann dich gut hören')

请问有人知道问题出在哪吗?恳请提供帮助。


解决方案

导致爆音和播放异常的核心问题有两个,修正后即可正常播放:

  • 采样率不匹配:你在sd.OutputStream中硬编码了采样率为24000,但实际获取的设备默认采样率是target_sr。如果两者不一致,播放时音频的速度会被错误拉伸,直接导致爆音和时长异常。必须将OutputStream的采样率参数改为target_sr。

  • 异步上下文同步IO阻塞风险:stream.write()是同步方法,在异步循环中直接调用可能导致阻塞,影响流式播放的连贯性。可以将写入操作包装在异步线程池中执行,避免阻塞事件循环。

修正后的代码如下:

import asyncio
import numpy as np
import sounddevice as sd
from scipy.signal import resample_poly
import edge_tts

EDGE_SR = 24000

async def stream_speak(text, voice="de-DE-SeraphinaMultilingualNeural"):
    communicate = edge_tts.Communicate(text, voice=voice)

    device_info = sd.query_devices(kind='output')
    target_sr = int(device_info['default_samplerate'])
    print(f"Playing at {target_sr} Hz (resampling from {EDGE_SR} Hz)")

    # 使用target_sr初始化OutputStream,匹配设备采样率
    with sd.OutputStream(samplerate=target_sr, blocksize=1024, channels=1, dtype="int16") as stream:
        async for chunk in communicate.stream():
            if chunk["type"] == "audio":
                audio = np.frombuffer(chunk["data"], dtype=np.int16)
                if target_sr != EDGE_SR:
                    # 重采样时确保计算正确,resample_poly的参数是(目标采样率, 原采样率)
                    audio = resample_poly(audio, target_sr, EDGE_SR).astype(np.int16)
                # 使用线程池执行同步写入,避免阻塞异步循环
                await asyncio.get_event_loop().run_in_executor(None, stream.write, audio)

def speak_sync(text):
    asyncio.run(stream_speak(text))

speak_sync('Hallo, ich kann dich gut hören')

额外注意事项:

  • 确保scipy库已正确安装(resample_poly依赖它),如果不想依赖scipy,也可以使用其他轻量重采样库如librosa的resample方法。
  • 部分设备可能对blocksize参数敏感,如果仍有异常,可以尝试调整blocksize为2048或4096,找到最适合设备的数值。

内容的提问来源于stack exchange,提问作者Artem Melnyk

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.12 09:22:11