FastAPI搭建Piper TTS服务器:流式WAV文件格式异常问题求助
解决FastAPI流式输出Piper TTS WAV的兼容性问题
方案一:修改WAV头的音频大小字段为最大值
WAV格式的data块确实需要提前声明数据大小,但流式场景下无法预知,你可以把这个字段设为32位无符号整数的最大值(0xFFFFFFFF)。绝大多数播放器和音频编辑器会忽略这个过大的值,直接读取到流结束,不会出现崩溃或无法播放的问题,完美解决FireFox和Audacity的兼容问题。
手动构造符合要求的WAV头(替代wave模块的默认生成方式):
import struct def build_streaming_wav_header(sample_rate: int, channels: int = 1, bits_per_sample: int = 16): # 计算WAV头各部分参数 audio_format = 1 # PCM格式 byte_rate = sample_rate * channels * (bits_per_sample // 8) block_align = channels * (bits_per_sample // 8) # 拼接RIFF块 riff_header = b"RIFF" + struct.pack('<I', 0xFFFFFFFF) + b"WAVE" # 拼接fmt块 fmt_chunk = ( b"fmt " + struct.pack('<I', 16) + struct.pack('<H', audio_format) + struct.pack('<H', channels) + struct.pack('<I', sample_rate) + struct.pack('<I', byte_rate) + struct.pack('<H', block_align) + struct.pack('<H', bits_per_sample) ) # 拼接data块(大小设为最大值) data_chunk = b"data" + struct.pack('<I', 0xFFFFFFFF) return riff_header + fmt_chunk + data_chunk # 在接口中使用 @app.get("/stream-tts-wav") async def stream_tts_wav(text: str): # 生成流式WAV头 wav_header = build_streaming_wav_header(model_metadata.sample_rate) yield wav_header # 流式输出音频数据 for chunk in synthesize_stream_raw(text, ...): yield chunk
方案二:更换为流式友好的音频格式(OGG Opus)
如果不想纠结WAV的格式限制,换用OGG Opus是更优的选择:
- 天生支持流式传输,无需提前知道总数据大小
- 所有主流浏览器(包括FireFox)、音频编辑器(Audacity)完美兼容
- 相同音质下文件体积比WAV小很多
可以用FFmpeg将Piper输出的原始PCM实时编码为OGG Opus:
import subprocess from fastapi import Response @app.get("/stream-tts-opus") async def stream_tts_opus(text: str): sample_rate = model_metadata.sample_rate # 启动FFmpeg进程,实时编码PCM为OGG Opus ffmpeg_cmd = [ "ffmpeg", "-f", "s16le", "-ar", str(sample_rate), "-ac", "1", "-i", "-", "-c:a", "libopus", "-f", "ogg", "-" ] process = subprocess.Popen( ffmpeg_cmd, stdin=subprocess.PIPE, stdout=subprocess.PIPE, stderr=subprocess.DEVNULL ) def stream_generator(): try: # 把Piper的PCM流喂给FFmpeg for pcm_chunk in synthesize_stream_raw(text, ...): process.stdin.write(pcm_chunk) process.stdin.flush() # 读取编码后的OGG数据并输出 while output_chunk := process.stdout.read(4096): yield output_chunk finally: process.stdin.close() process.wait() return Response(stream_generator(), media_type="audio/ogg")
总结
- 若坚持使用WAV格式,方案一的修改成本最低,兼容性足够覆盖绝大多数场景
- 若追求更好的兼容性和传输效率,方案二的OGG Opus是长期最优解
内容的提问来源于stack exchange,提问作者Bill Trần
相关产品推荐
相关产品推荐

