You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

拼接本地MP3与API返回音频后前端仅播放开头的问题排查

MP3文件与API返回音频拼接后无法完整播放问题排查与解决

问题场景

我有一个MP3文件,调用返回音频的API获取音频内容,想把本地MP3和API返回的音频拼接后,将二进制数据返回给前端。写了如下FastAPI接口:

@app.get("/audio")
async def generate_audio():
    summaries = await get_summaries()
    audio_content = []
    for summary in summaries:
        audio = await text_to_speech(summary)
        if audio.status_code == 200:
            audio_content.append(audio.content)
    
    with open("news_jingle.mp3", "rb") as f:
        jingle = f.read()
    
    content = jingle + b''.join(audio_content)
    return Response(content=content, media_type="audio/mpeg")

但前端播放时,只播放开头的MP3铃声就停止,后续音频完全不播放。单独返回铃声或单独返回API获取的音频都能正常播放,尝试过pydub和ffmpeg但没成功。

问题根源

直接拼接MP3的二进制内容行不通,因为MP3文件不是纯流式音频,每个MP3文件都包含自己的文件头、帧结构和元数据。播放器解析时只会识别第一个MP3的文件头,播放完第一个文件的音频帧后,就会认为音频结束,不会处理后面拼接的二进制内容。

正确解决方案

方法1:使用pydub正确拼接(需安装ffmpeg)

先确保pydub和ffmpeg环境配置正确:

  • 安装pydub:pip install pydub
  • 安装ffmpeg并配置到系统环境变量(或在代码中指定路径)

修改后的FastAPI代码:

from pydub import AudioSegment
from fastapi import Response
import io

@app.get("/audio")
async def generate_audio():
    summaries = await get_summaries()
    
    # 加载本地铃声
    jingle = AudioSegment.from_mp3("news_jingle.mp3")
    combined_audio = jingle
    
    for summary in summaries:
        audio = await text_to_speech(summary)
        if audio.status_code == 200:
            # 从API返回的二进制内容加载音频
            api_audio = AudioSegment.from_file(io.BytesIO(audio.content), format="mp3")
            # 拼接音频
            combined_audio += api_audio
    
    # 将拼接后的音频转为MP3二进制
    output = io.BytesIO()
    combined_audio.export(output, format="mp3")
    output.seek(0)
    
    return Response(content=output.read(), media_type="audio/mpeg")

方法2:使用ffmpeg命令行拼接(适合批量或复杂场景)

如果pydub调用有问题,可以直接用ffmpeg命令行生成拼接后的音频,再返回给前端。示例代码:

import subprocess
import io
from fastapi import Response
import os

@app.get("/audio")
async def generate_audio():
    summaries = await get_summaries()
    # 保存API返回的音频到临时文件
    temp_files = ["news_jingle.mp3"]
    for idx, summary in enumerate(summaries):
        audio = await text_to_speech(summary)
        if audio.status_code == 200:
            temp_path = f"temp_audio_{idx}.mp3"
            with open(temp_path, "wb") as f:
                f.write(audio.content)
            temp_files.append(temp_path)
    
    # 生成ffmpeg拼接指令文件
    concat_file_path = "concat_list.txt"
    with open(concat_file_path, "w") as f:
        for file in temp_files:
            f.write(f"file '{file}'\n")
    
    # 执行ffmpeg拼接,输出到内存流
    output = io.BytesIO()
    subprocess.run(
        [
            "ffmpeg", "-f", "concat", "-safe", "0",
            "-i", concat_file_path, "-c", "copy", "-f", "mp3", "-"
        ],
        stdout=output,
        stderr=subprocess.PIPE,
        check=True
    )
    output.seek(0)
    
    # 清理临时文件
    for file in temp_files[1:]:
        os.remove(file)
    os.remove(concat_file_path)
    
    return Response(content=output.read(), media_type="audio/mpeg")

注意事项

  • 之前用pydub失败大概率是ffmpeg未正确配置,确保ffmpeg可执行文件在系统PATH中,或在代码中指定路径:AudioSegment.converter = "path/to/ffmpeg"
  • 拼接时尽量保证所有音频的采样率、声道数、比特率一致,pydub会自动处理格式转换;ffmpeg的-c copy参数要求格式完全一致,否则需添加转码参数(如-c:a libmp3lame)。

内容的提问来源于stack exchange,提问作者Ammar Ibraheem

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.25 03:35:00