拼接本地MP3与API返回音频后前端仅播放开头的问题排查
MP3文件与API返回音频拼接后无法完整播放问题排查与解决
问题场景
我有一个MP3文件,调用返回音频的API获取音频内容,想把本地MP3和API返回的音频拼接后,将二进制数据返回给前端。写了如下FastAPI接口:
@app.get("/audio") async def generate_audio(): summaries = await get_summaries() audio_content = [] for summary in summaries: audio = await text_to_speech(summary) if audio.status_code == 200: audio_content.append(audio.content) with open("news_jingle.mp3", "rb") as f: jingle = f.read() content = jingle + b''.join(audio_content) return Response(content=content, media_type="audio/mpeg")
但前端播放时,只播放开头的MP3铃声就停止,后续音频完全不播放。单独返回铃声或单独返回API获取的音频都能正常播放,尝试过pydub和ffmpeg但没成功。
问题根源
直接拼接MP3的二进制内容行不通,因为MP3文件不是纯流式音频,每个MP3文件都包含自己的文件头、帧结构和元数据。播放器解析时只会识别第一个MP3的文件头,播放完第一个文件的音频帧后,就会认为音频结束,不会处理后面拼接的二进制内容。
正确解决方案
方法1:使用pydub正确拼接(需安装ffmpeg)
先确保pydub和ffmpeg环境配置正确:
- 安装pydub:
pip install pydub - 安装ffmpeg并配置到系统环境变量(或在代码中指定路径)
修改后的FastAPI代码:
from pydub import AudioSegment from fastapi import Response import io @app.get("/audio") async def generate_audio(): summaries = await get_summaries() # 加载本地铃声 jingle = AudioSegment.from_mp3("news_jingle.mp3") combined_audio = jingle for summary in summaries: audio = await text_to_speech(summary) if audio.status_code == 200: # 从API返回的二进制内容加载音频 api_audio = AudioSegment.from_file(io.BytesIO(audio.content), format="mp3") # 拼接音频 combined_audio += api_audio # 将拼接后的音频转为MP3二进制 output = io.BytesIO() combined_audio.export(output, format="mp3") output.seek(0) return Response(content=output.read(), media_type="audio/mpeg")
方法2:使用ffmpeg命令行拼接(适合批量或复杂场景)
如果pydub调用有问题,可以直接用ffmpeg命令行生成拼接后的音频,再返回给前端。示例代码:
import subprocess import io from fastapi import Response import os @app.get("/audio") async def generate_audio(): summaries = await get_summaries() # 保存API返回的音频到临时文件 temp_files = ["news_jingle.mp3"] for idx, summary in enumerate(summaries): audio = await text_to_speech(summary) if audio.status_code == 200: temp_path = f"temp_audio_{idx}.mp3" with open(temp_path, "wb") as f: f.write(audio.content) temp_files.append(temp_path) # 生成ffmpeg拼接指令文件 concat_file_path = "concat_list.txt" with open(concat_file_path, "w") as f: for file in temp_files: f.write(f"file '{file}'\n") # 执行ffmpeg拼接,输出到内存流 output = io.BytesIO() subprocess.run( [ "ffmpeg", "-f", "concat", "-safe", "0", "-i", concat_file_path, "-c", "copy", "-f", "mp3", "-" ], stdout=output, stderr=subprocess.PIPE, check=True ) output.seek(0) # 清理临时文件 for file in temp_files[1:]: os.remove(file) os.remove(concat_file_path) return Response(content=output.read(), media_type="audio/mpeg")
注意事项
- 之前用pydub失败大概率是ffmpeg未正确配置,确保ffmpeg可执行文件在系统PATH中,或在代码中指定路径:
AudioSegment.converter = "path/to/ffmpeg" - 拼接时尽量保证所有音频的采样率、声道数、比特率一致,pydub会自动处理格式转换;ffmpeg的
-c copy参数要求格式完全一致,否则需添加转码参数(如-c:a libmp3lame)。
内容的提问来源于stack exchange,提问作者Ammar Ibraheem
相关产品推荐
相关产品推荐

