You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

前端音频Blob传FastAPI后端合并WAV遇Libsndfile格式识别错误

问题:前端录制音频通过SocketIO传FastAPI后端合并时触发格式不识别错误

问题说明

将前端录制的音频Blob通过SocketIO发送到FastAPI后端,实现WAV文件合并,但测试时触发soundfile.LibsndfileError,提示“Format not recognised”,文件虽能保存但无法正常打开。

后端FastAPI SocketIO代码

@sio.on('voice')
def get_chunk(sid, chunk: bytes):
    directory = str("dddd")
    if not os.path.exists(directory):
        os.makedirs(directory)
    file_path = os.path.join(directory, f'{sid}.wav')
    file_path_chunk =  os.path.join(directory, f'{sid}_chunk.wav')
    outfile =  os.path.join(directory, f'{sid}_result.wav')
    if not os.path.exists(file_path):  # 首次传输无需合并
        with open(file_path, 'wb+') as file:
            file.write(chunk)
    else:
        with open(file_path_chunk, 'wb+') as file:
            file.write(chunk)
        try:
            # 读取原文件
            data, samplerate = soundfile.read(file_path)
            # 读取新分片文件
            chunk_data, _ = soundfile.read(file_path_chunk)
            # 合并数据
            data = np.concatenate([data, chunk_data])
            # 写入结果文件
            soundfile.write(outfile, data, samplerate)
        finally:
            # 删除分片文件
            os.remove(file_path_chunk)

前端JavaScript代码

<!DOCTYPE html>
<html lang="en">
<head>
    <meta charset="UTF-8">
    <title>마이크 테스트</title>
    <script src="https://cdnjs.cloudflare.com/ajax/libs/socket.io/4.5.2/socket.io.js"></script>
</head>
<body>
    <input type=checkbox id="chk-hear-mic"><label for="chk-hear-mic">마이크 소리 듣기</label>
    <button id="record">녹음</button>
    <button id="stop">정지</button>
    <div id="sound-clips"></div>
    <script>
        var socket = io('http://127.0.0.1:5000');
        const record = document.getElementById("record")
        const stop = document.getElementById("stop")
        const soundClips = document.getElementById("sound-clips")
        const chkHearMic = document.getElementById("chk-hear-mic")
        const audioCtx = new(window.AudioContext || window.webkitAudioContext)()
        const analyser = audioCtx.createAnalyser()
        function makeSound(stream) {
            const source = audioCtx.createMediaStreamSource(stream)
            socket.connect()
            source.connect(analyser)
            analyser.connect(audioCtx.destination)
        }
        if (navigator.mediaDevices) {
            console.log('getUserMedia supported.')
            const constraints = { audio: true }
            let chunks = []
            navigator.mediaDevices.getUserMedia(constraints)
                .then(stream => {
                    const mediaRecorder = new MediaRecorder(stream)
                    chkHearMic.onchange = e => {
                        if(e.target.checked == true) {
                            audioCtx.resume()
                            makeSound(stream)
                        } else {
                            audioCtx.suspend()
                        }
                    }
                    record.onclick = () => {
                        mediaRecorder.start()
                        console.log(mediaRecorder.state)
                        console.log("recorder started")
                        record.style.background = "red"
                        record.style.color = "black"
                    }
                    stop.onclick = () => {
                        mediaRecorder.stop()
                        console.log(mediaRecorder.state)
                        console.log("recorder stopped")
                        record.style.background = ""
                        record.style.color = ""
                    }
                    mediaRecorder.onstop = e => {
                        console.log("data available after MediaRecorder.stop() called.")
                        const clipName = prompt("오디오 파일 제목을 입력하세요.", new Date())
                        const clipContainer = document.createElement('article')
                        const clipLabel = document.createElement('p')
                        const audio = document.createElement('audio')
                        const deleteButton = document.createElement('button')
                        clipContainer.classList.add('clip')
                        audio.setAttribute('controls', '')
                        deleteButton.innerHTML = "삭제"
                        clipLabel.innerHTML = clipName
                        clipContainer.appendChild(audio)
                        clipContainer.appendChild(clipLabel)
                        clipContainer.appendChild(deleteButton)
                        soundClips.appendChild(clipContainer)
                        audio.controls = true
                        const blob = new Blob(chunks, {
                            'type': 'audio/ogg codecs=opus'
                        })
                        chunks = []
                        const audioURL = URL.createObjectURL(blob)
                        audio.src = audioURL
                        console.log("recorder stopped")
                        deleteButton.onclick = e => {
                            evtTgt = e.target
                            evtTgt.parentNode.parentNode.removeChild(evtTgt.parentNode)
                        }
                    }
                })
                .catch(err => {
                    console.log('The following error occurred: ' + err)
                })
        }
    </script>
</body></html>

报错信息

Task exception was never retrieved
future: <Task finished name='Task-29' coro=<InstrumentedAsyncServer._handle_event_internal() done, defined at F:\fastapi-socketio-wb38\.vent\Lib\site-packages\socketio\async_admin.py:274> exception=LibsndfileError(1, "Error opening 'dddd\\-dZWDNG1oD_HKvhjAAAB.wav': ")>
Traceback (most recent call last):
  File "F:\fastapi-socketio-wb38\.vent\Lib\site-packages\socketio\async_admin.py", line 276, in _handle_event_internal
    ret = await self.sio.__handle_event_internal(server, sid, eio_sid,
          ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "F:\fastapi-socketio-wb38\.vent\Lib\site-packages\socketio\async_server.py", line 597, in _handle_event_internal
    r = await server._trigger_event(data[0], namespace, sid, *data[1:])
        ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "F:\fastapi-socketio-wb38\.vent\Lib\site-packages\socketio\async_server.py", line 635, in _trigger_event
    ret = handler(*args)
          ^^^^^^^^^^^^^^
  File "f:\fastapi-socketio-wb38\Python-Javascript-Websocket-Video-Streaming--main\ppom.py", line 160, in get_chunk
    data, samplerate = soundfile.read(file_path)
                       ^^^^^^^^^^^^^^^^^^^^^^^^^
  File "F:\fastapi-socketio-wb38\.vent\Lib\site-packages\soundfile.py", line 285, in read
    with SoundFile(file, 'r', samplerate, channels,
         ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "F:\fastapi-socketio-wb38\.vent\Lib\site-packages\soundfile.py", line 658, in __init__
    self._file = self._open(file, mode_int, closefd)
                 ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "F:\fastapi-socketio-wb38\.vent\Lib\site-packages\soundfile.py", line 1216, in _open
    raise LibsndfileError(err, prefix="Error opening {0!r}: ".format(self.name))
soundfile.LibsndfileError: Error opening 'dddd\-dZWDNG1oD_HKvhjAAAB.wav': Format not recognised.
Task exception was never retrieved
future: <Task finished name='Task-32' coro=<InstrumentedAsyncServer._handle_event_internal() done, defined at F:\fastapi-socketio-wb38\.vent\Lib\site-packages\socketio\async_admin.py:274> exception=LibsndfileError(1, "Error opening 'dddd\\-dZWDNG1oD_HKvhjAAAB_chunk.wav': ")>
Traceback (most recent call last):
  File "F:\fastapi-socketio-wb38\.vent\Lib\site-packages\socketio\async_admin.py", line 276, in _handle_event_internal
    ret = await self.sio.__handle_event_internal(server, sid, eio_sid,
          ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "F:\fastapi-socketio-wb38\.vent\Lib\site-packages\socketio\async_server.py", line 597, in _handle_event_internal
    r = await server._trigger_event(data[0], namespace, sid, *data[1:])
        ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "F:\fastapi-socketio-wb38\.vent\Lib\site-packages\socketio\async_server.py", line 635, in _trigger_event
    ret = handler(*args)
          ^^^^^^^^^^^^^^
  File "f:\fastapi-socketio-wb38\Python-Javascript-Websocket-Video-Streaming--main\ppom.py", line 166, in get_chunk
    data, samplerate = soundfile.read(file_path_chunk)
                       ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "F:\fastapi-socketio-wb38\.vent\Lib\site-packages\soundfile.py", line 285, in read
    with SoundFile(file, 'r', samplerate, channels,
         ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "F:\fastapi-socketio-wb38\.vent\Lib\site-packages\soundfile.py", line 658, in __init__
    self._file = self._open(file, mode_int, closefd)
                 ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "F:\fastapi-socketio-wb38\.vent\Lib\site-packages\soundfile.py", line 1216, in _open
    raise LibsndfileError(err, prefix="Error opening {0!r}: ".format(self.name))
soundfile.LibsndfileError: Error opening 'dddd\-dZWDNG1oD_HKvhjAAAB_chunk.wav': Format not recognised.

解决方法

1. 前端补全音频数据发送逻辑

当前前端未监听MediaRecorder的ondataavailable事件,导致录制的音频数据未发送到后端。添加以下代码:

mediaRecorder.ondataavailable = e => {
    chunks.push(e.data);
    // 发送音频分片到后端
    socket.emit('voice', e.data);
};

同时可在调用mediaRecorder.start()时指定分片间隔,比如每1秒发送一次:

record.onclick = () => {
    mediaRecorder.start(1000); // 每1000ms生成一个分片
    // ... 原有代码
}

2. 解决格式不匹配问题

前端录制的是OGG/Opus格式音频,但后端直接将其保存为.wav后缀,soundfile库无法识别该格式。需先将OGG转成WAV再处理,推荐用pydub库实现格式转换:

步骤1:安装依赖

pip install pydub

同时需安装ffmpeg,并将其路径添加到系统环境变量中。

步骤2:修改后端代码

from pydub import AudioSegment
import os
import numpy as np
import soundfile as sf

@sio.on('voice')
def get_chunk(sid, chunk: bytes):
    directory = "dddd"
    if not os.path.exists(directory):
        os.makedirs(directory)
    
    # 临时保存前端发送的OGG数据
    temp_ogg = os.path.join(directory, f'{sid}_temp.ogg')
    with open(temp_ogg, 'wb+') as f:
        f.write(chunk)
    
    # 转换为WAV格式
    audio = AudioSegment.from_file(temp_ogg, format="ogg")
    main_wav = os.path.join(directory, f'{sid}.wav')
    chunk_wav = os.path.join(directory, f'{sid}_chunk.wav')
    result_wav = os.path.join(directory, f'{sid}_result.wav')
    
    if not os.path.exists(main_wav):
        # 首次传输,直接保存转换后的WAV
        audio.export(main_wav, format="wav")
    else:
        # 保存当前分片的WAV
        audio.export(chunk_wav, format="wav")
        try:
            # 读取并合并音频数据
            data, samplerate = sf.read(main_wav)
            chunk_data, _ = sf.read(chunk_wav)
            merged_data = np.concatenate([data, chunk_data])
            sf.write(result_wav, merged_data, samplerate)
            # 将合并后的文件替换原主文件
            os.replace(result_wav, main_wav)
        finally:
            os.remove(chunk_wav)
    
    # 删除临时OGG文件
    os.remove(temp_ogg)

内容的提问来源于stack exchange,提问作者a_crszkvc30Last_NameCol

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.02 03:37:03