You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Flask后端无法识别音频Blob,Librosa/Pydub处理报错求助

问题描述

将前端录制的音频Blob发送到Flask后端,尝试用librosa/pydub转换为梅尔频谱图时出现以下错误:

ERROR:root:Librosa error processing <_io.BytesIO object at 0x3219eff60>: Error opening <_io.BytesIO object at 0x3219eff60>: Format not recognised.
ERROR:root:Pydub error processing <_io.BytesIO object at 0x3219eff60>: Decoding failed. ffmpeg returned error code: 183

前端使用语音录制库,通过以下代码打包Blob并发送:

const onAudioDownload = (blob) => {
    console.log('audio ended')
    console.log(typeof(blob))

    const formData = new FormData();
    formData.append('file', new Blob([blob], { type: 'audio/wav' }));

    fetch('http://127.0.0.1:5000/predictTest', {
      method: 'POST',
      body: formData,
    })
      .then((response) => {
        if (!response.ok) {
          throw new Error('Network response was not ok');
        }
        return response.json();
      })
      .then((data) => {
        console.log('Success:', data);
      })
      .catch((error) => {
        console.error('Error:', error);
      });
  };

后端处理代码:

def process_audio(song):

    sr=16000
    n_mels=128
    n_fft=2048
    hop_length=512

    slice_length = 911

    print(song)
    print('Loading song...')
    try:
        y, sr = librosa.load(song, sr=sr)
    except Exception as e:
        logging.error(f"Librosa error processing {song}: {e}")
        # Fallback: Use pydub to decode the MP3 file
        try:
            audio = AudioSegment.from_file(song)
            y = np.array(audio.get_array_of_samples())
            sr = audio.frame_rate
        except Exception as e:
            logging.error(f"Pydub error processing {song}: {e}")
解决方案

1. 前端修正Blob包装逻辑

该语音录制库默认生成MP3格式Blob,强行包装为audio/wav类型会导致实际编码与声明格式不匹配,后端无法识别。直接使用原始Blob即可:

const onAudioDownload = (blob) => {
    console.log('audio ended')
    console.log(blob.type) // 可打印确认原始格式,通常为audio/mp3

    const formData = new FormData();
    // 直接传入原始Blob,无需重新包装
    formData.append('file', blob, 'recording.mp3');

    fetch('http://127.0.0.1:5000/predictTest', {
      method: 'POST',
      body: formData,
    })
      .then((response) => {
        if (!response.ok) {
          throw new Error('Network response was not ok');
        }
        return response.json();
      })
      .then((data) => {
        console.log('Success:', data);
      })
      .catch((error) => {
        console.error('Error:', error);
      });
  };

2. 后端优化音频处理流程

Librosa和pydub直接处理BytesIO对象易出现格式识别问题,建议先将上传文件保存为临时文件,再明确指定格式进行处理:

import tempfile
import librosa
import numpy as np
from pydub import AudioSegment
import logging
import os

def process_audio(song_file):
    sr=16000
    n_mels=128
    n_fft=2048
    hop_length=512
    slice_length = 911

    print('Loading song...')
    # 创建临时文件存储上传音频
    with tempfile.NamedTemporaryFile(suffix='.mp3', delete=False) as temp_file:
        temp_file.write(song_file.read())
        temp_path = temp_file.name

    try:
        # Librosa加载临时文件
        y, sr = librosa.load(temp_path, sr=sr)
        # 此处添加梅尔频谱图生成逻辑
    except Exception as e:
        logging.error(f"Librosa error processing {temp_path}: {e}")
        try:
            # 明确指定音频格式为mp3,避免自动识别错误
            audio = AudioSegment.from_file(temp_path, format='mp3')
            # 转换为单声道(若原始为立体声)
            if audio.channels > 1:
                audio = audio.set_channels(1)
            y = np.array(audio.get_array_of_samples())
            sr = audio.frame_rate
            # 此处添加梅尔频谱图生成逻辑
        except Exception as e:
            logging.error(f"Pydub error processing {temp_path}: {e}")
    finally:
        # 清理临时文件
        os.unlink(temp_path)

3. 确保FFmpeg正确配置

Pydub依赖FFmpeg,错误码183多为FFmpeg未找到或版本不兼容:

  • 下载安装FFmpeg,将其路径添加到系统环境变量PATH中
  • 若无法修改系统环境变量,可在代码中直接指定FFmpeg路径:
from pydub import AudioSegment
# Windows系统示例
AudioSegment.converter = "C:/ffmpeg/bin/ffmpeg.exe"
# Linux/Mac系统示例
AudioSegment.converter = "/usr/bin/ffmpeg"

内容的提问来源于stack exchange,提问作者Lan Do

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.20 18:13:11