You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python Sounddevice播放音频速度过快的问题求助

问题:sounddevice播放音频速度过快,无法听清内容

我用Python的sounddevice模块播放音频流或本地WAV文件时,扬声器输出的速度快到完全听不清,但用VLC播放同一WAV文件却一切正常。试了两段代码都没解决问题,代码如下:

第一段代码(sd.play方式)

import sounddevice as sd
import soundfile as sf
import wave
from piper.voice import PiperVoice

sd.default.device = 1
filename = 'temp.wav'

voicedir = "./piper/" # 本地onnx模型存放路径
model = voicedir+"en_GB-alan-low.onnx"
voice = PiperVoice.load(model)
wav_file = wave.open(filename, 'w')
text = "This is an example of text-to-speech using Piper TTS."
audio = voice.synthesize(text,wav_file)

# 读取文件数据和采样率
data, fs = sf.read(filename)
sd.play(data, fs)
status = sd.wait()  # 等待播放结束

第二段代码(流播放方式)

import sounddevice as sd
from piper.voice import PiperVoice
import numpy as np

voicedir = "./piper/" # 本地onnx模型存放路径
model = voicedir+"en_GB-alan-low.onnx"
voice = PiperVoice.load(model)
text = "This is an example of text-to-speech using Piper TTS."

# 配置sounddevice输出流参数
sd.default.device = 1
sd.default.samplerate = 16000
samplerate = sd.query_devices(device=1)
stream = sd.RawOutputStream(channels=1, dtype='int16')
print(voice.config.sample_rate)

stream.start()
for audio_bytes in voice.synthesize_stream_raw(text):
    int_data = np.frombuffer(audio_bytes, dtype=np.int16)
    stream.write(int_data)
stream.stop()
stream.Close()

问题根源与修复方案

第一段代码的问题与修复

问题出在WAV文件头信息未正确设置:用wave.open创建文件时,没有指定采样率、声道数、位深,导致生成的WAV文件头不完整,sf.read读取时解析出错误的采样率,最终播放速度异常。

修复后的代码:

import sounddevice as sd
import soundfile as sf
import wave
from piper.voice import PiperVoice

sd.default.device = 1
filename = 'temp.wav'

voicedir = "./piper/"
model = voicedir+"en_GB-alan-low.onnx"
voice = PiperVoice.load(model)

# 从Piper模型配置中获取正确的音频参数
sample_rate = voice.config.sample_rate
channels = voice.config.num_channels
sample_width = 2  # Piper生成16位PCM,对应2字节

# 初始化WAV文件时指定完整参数
wav_file = wave.open(filename, 'w')
wav_file.setnchannels(channels)
wav_file.setsampwidth(sample_width)
wav_file.setframerate(sample_rate)

text = "This is an example of text-to-speech using Piper TTS."
voice.synthesize(text, wav_file)
wav_file.close()  # 必须关闭文件,确保头信息写入完成

# 读取并播放
data, fs = sf.read(filename)
sd.play(data, fs)
status = sd.wait()

第二段代码的问题与修复

问题有三个:

  1. 手动设置的sd.default.samplerate = 16000和Piper模型的实际采样率不匹配(比如en_GB-alan-low.onnx的采样率是22050)
  2. 初始化RawOutputStream时未指定采样率,默认用了16000,导致播放采样率和音频数据不匹配
  3. 最后调用了错误的stream.Close()(应该是小写的stream.close())

修复后的代码:

import sounddevice as sd
from piper.voice import PiperVoice
import numpy as np

voicedir = "./piper/"
model = voicedir+"en_GB-alan-low.onnx"
voice = PiperVoice.load(model)
text = "This is an example of text-to-speech using Piper TTS."

sd.default.device = 1
# 直接使用Piper模型的采样率
sample_rate = voice.config.sample_rate
channels = voice.config.num_channels

# 初始化流时指定完整的正确参数
stream = sd.RawOutputStream(
    samplerate=sample_rate,
    channels=channels,
    dtype='int16'
)

print(f"模型采样率:{sample_rate}")

stream.start()
for audio_bytes in voice.synthesize_stream_raw(text):
    int_data = np.frombuffer(audio_bytes, dtype=np.int16)
    stream.write(int_data)
stream.stop()
stream.close()

通用排查要点

  • 始终确保sounddevice的播放采样率和Piper模型的voice.config.sample_rate完全一致
  • Piper生成的是16位有符号整数PCM数据,所以播放时dtype必须设为'int16'
  • 生成WAV文件时必须完整设置头信息,否则播放器无法正确解析音频参数

内容的提问来源于stack exchange,提问作者stefan van leeuwen

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.22 02:17:06