使用Python Sounddevice播放音频速度过快的问题求助
问题:sounddevice播放音频速度过快,无法听清内容
我用Python的sounddevice模块播放音频流或本地WAV文件时,扬声器输出的速度快到完全听不清,但用VLC播放同一WAV文件却一切正常。试了两段代码都没解决问题,代码如下:
第一段代码(sd.play方式)
import sounddevice as sd import soundfile as sf import wave from piper.voice import PiperVoice sd.default.device = 1 filename = 'temp.wav' voicedir = "./piper/" # 本地onnx模型存放路径 model = voicedir+"en_GB-alan-low.onnx" voice = PiperVoice.load(model) wav_file = wave.open(filename, 'w') text = "This is an example of text-to-speech using Piper TTS." audio = voice.synthesize(text,wav_file) # 读取文件数据和采样率 data, fs = sf.read(filename) sd.play(data, fs) status = sd.wait() # 等待播放结束
第二段代码(流播放方式)
import sounddevice as sd from piper.voice import PiperVoice import numpy as np voicedir = "./piper/" # 本地onnx模型存放路径 model = voicedir+"en_GB-alan-low.onnx" voice = PiperVoice.load(model) text = "This is an example of text-to-speech using Piper TTS." # 配置sounddevice输出流参数 sd.default.device = 1 sd.default.samplerate = 16000 samplerate = sd.query_devices(device=1) stream = sd.RawOutputStream(channels=1, dtype='int16') print(voice.config.sample_rate) stream.start() for audio_bytes in voice.synthesize_stream_raw(text): int_data = np.frombuffer(audio_bytes, dtype=np.int16) stream.write(int_data) stream.stop() stream.Close()
问题根源与修复方案
第一段代码的问题与修复
问题出在WAV文件头信息未正确设置:用wave.open创建文件时,没有指定采样率、声道数、位深,导致生成的WAV文件头不完整,sf.read读取时解析出错误的采样率,最终播放速度异常。
修复后的代码:
import sounddevice as sd import soundfile as sf import wave from piper.voice import PiperVoice sd.default.device = 1 filename = 'temp.wav' voicedir = "./piper/" model = voicedir+"en_GB-alan-low.onnx" voice = PiperVoice.load(model) # 从Piper模型配置中获取正确的音频参数 sample_rate = voice.config.sample_rate channels = voice.config.num_channels sample_width = 2 # Piper生成16位PCM,对应2字节 # 初始化WAV文件时指定完整参数 wav_file = wave.open(filename, 'w') wav_file.setnchannels(channels) wav_file.setsampwidth(sample_width) wav_file.setframerate(sample_rate) text = "This is an example of text-to-speech using Piper TTS." voice.synthesize(text, wav_file) wav_file.close() # 必须关闭文件,确保头信息写入完成 # 读取并播放 data, fs = sf.read(filename) sd.play(data, fs) status = sd.wait()
第二段代码的问题与修复
问题有三个:
- 手动设置的
sd.default.samplerate = 16000和Piper模型的实际采样率不匹配(比如en_GB-alan-low.onnx的采样率是22050) - 初始化
RawOutputStream时未指定采样率,默认用了16000,导致播放采样率和音频数据不匹配 - 最后调用了错误的
stream.Close()(应该是小写的stream.close())
修复后的代码:
import sounddevice as sd from piper.voice import PiperVoice import numpy as np voicedir = "./piper/" model = voicedir+"en_GB-alan-low.onnx" voice = PiperVoice.load(model) text = "This is an example of text-to-speech using Piper TTS." sd.default.device = 1 # 直接使用Piper模型的采样率 sample_rate = voice.config.sample_rate channels = voice.config.num_channels # 初始化流时指定完整的正确参数 stream = sd.RawOutputStream( samplerate=sample_rate, channels=channels, dtype='int16' ) print(f"模型采样率:{sample_rate}") stream.start() for audio_bytes in voice.synthesize_stream_raw(text): int_data = np.frombuffer(audio_bytes, dtype=np.int16) stream.write(int_data) stream.stop() stream.close()
通用排查要点
- 始终确保sounddevice的播放采样率和Piper模型的
voice.config.sample_rate完全一致 - Piper生成的是16位有符号整数PCM数据,所以播放时dtype必须设为
'int16' - 生成WAV文件时必须完整设置头信息,否则播放器无法正确解析音频参数
内容的提问来源于stack exchange,提问作者stefan van leeuwen
相关产品推荐
相关产品推荐

