You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Django语音转文本视图报错:音频无法被读取为PCM WAV等支持格式

问题复现

我的语音转文本视图无法正常运行,已有wav格式文件,但运行时报出如下错误:

Traceback (most recent call last):
  File "C:\Users\privet01\Desktop\python projects\projex x\myvev\lib\site-packages\django\core\handlers\exception.py", line 47, in inner
    response = get_response(request)
  File "C:\Users\privet01\Desktop\python projects\projex x\myvev\lib\site-packages\django\core\handlers\base.py", line 179, in _get_response
    response = wrapped_callback(request, *callback_args, **callback_kwargs)
  File "C:\Users\privet01\Desktop\python projects\projex x\alpha1\appsolve\views.py", line 636, in spechtotext
    with sr.AudioFile(sound) as source:
  File "C:\Users\privet01\Desktop\python projects\projex x\myvev\lib\site-packages\speech_recognition\__init__.py", line 236, in __enter__
    raise ValueError("Audio file could not be read as PCM WAV, AIFF/AIFF-C, or Native FLAC; check if file is corrupted or in another format")
ValueError: Audio file could not be read as PCM WAV, AIFF/AIFF-C, or Native FLAC; check if file is corrupted or in another format
[23/Aug/2021 13:26:44] "POST /spechtotext/ HTTP/1.1" 500 152359

原视图代码

import speech_recognition as sr
from django.core.files.storage import FileSystemStorage

def spechtotext(request)   :
    if request.method=='POST':
        if request.FILES['theFileINEEd']  and '.mp3' in  str(request.FILES['theFileINEEd'].name).lower() :
            myfile = request.FILES['theFileINEEd']
            fs = FileSystemStorage('')
            filename = fs.save(str(request.user)+"speachtotext.WAV", myfile)
            file_url = fs.url(filename)
            full_link='http://127.0.0.1:8000'+file_url # changeble to life mode 
            r = sr.Recognizer()
            print(type(filename))
            print('hhhhhhhhhhhhhhhhhhhhhhhhhhhhhhhhhhhhhhhhhhhhhhhhhhhhhhhhhh')
            r = sr.Recognizer()
            import sys
            print(sys.version_info)
            print('qqqqqqqqqqqqqqqqqqqqqqqqq'+str(filename)+'aaaaaaaaaaaaaaaaaaaa')
            hellow=sr.AudioFile(str(filename))
            sound = str(filename)
            with sr.AudioFile(sound) as source:
                r.adjust_for_ambient_noise(source)
                print("Converting Audio To Text ..... ")
                audio = r.listen(source)
            try:
                print("Converted Audio Is : \n" + r.recognize_google(audio))
            except Exception as e:
                print("Error {} : ".format(e) )
                audio_text = r.listen(source)
                try:
                    # using google speech recognition
                    text = r.recognize_google(audio_text)
                    print('Converting audio transcripts into text ...')
                    print("this is the text : ",text)
                except:
                    print('Sorry.. run again...')   
            return render (request,'appsolve/spechtotext.html',{'l5arij':request.FILES['theFileINEEd'],'link':file_url,'fulllink':full_link})   
        else:
            return render (request,'appsolve/spechtotext.html',{'problem':"you must upload a file mp3"}) 
    else:
        return render (request,'appsolve/spechtotext.html',)  

相关材料

  • 网页报错截图:网页报错截图
  • 文件路径截图:文件路径截图

已尝试操作

  • 确认音频文件已存储在对应路径下
  • 分别测试了mp3和PCM格式的音频
  • 使用音质清晰的音频测试,问题仍未解决
解决方案

根因说明

你代码中直接将上传的mp3文件改后缀保存为.WAV格式,仅修改后缀不会改变文件本身的编码结构,speech_recognition库的AudioFile仅支持原生PCM编码的WAV、AIFF、FLAC格式,无法识别改了后缀的mp3文件,因此抛出对应错误。另外你代码中except块里调用r.listen(source)时,source已经被with块关闭,会额外触发报错。

修复步骤

  1. 安装依赖库
    你需要用pydub完成音频转码,先安装相关依赖:
pip install pydub

windows环境需要额外下载ffmpeg工具,将其bin目录添加到系统环境变量Path中,或在代码中指定ffmpeg路径。

  1. 修改视图代码
    核心修改音频存储和转码逻辑,修复异常块的错误逻辑:
import speech_recognition as sr
from django.core.files.storage import FileSystemStorage
from pydub import AudioSegment
import os

def spechtotext(request):
    if request.method=='POST':
        if request.FILES['theFileINEEd']  and '.mp3' in  str(request.FILES['theFileINEEd'].name).lower() :
            myfile = request.FILES['theFileINEEd']
            fs = FileSystemStorage('')
            # 先保存临时mp3文件
            temp_mp3_name = f"{str(request.user)}_temp.mp3"
            temp_mp3_path = fs.save(temp_mp3_name, myfile)
            # 转码为PCM编码的WAV文件,转成单声道16k采样率适配语音识别
            audio = AudioSegment.from_mp3(temp_mp3_path)
            wav_name = f"{str(request.user)}_speachtotext.WAV"
            audio.export(wav_name, format="wav", parameters=["-ac", "1", "-ar", "16000"])
            # 删除临时mp3文件
            os.remove(temp_mp3_path)
            file_url = fs.url(wav_name)
            full_link='http://127.0.0.1:8000'+file_url

            r = sr.Recognizer()
            with sr.AudioFile(wav_name) as source:
                r.adjust_for_ambient_noise(source)
                print("Converting Audio To Text ..... ")
                audio = r.listen(source)
            try:
                text = r.recognize_google(audio, language="zh-CN") # 可指定识别语言,默认是英文
                print("Converted Audio Is : \n" + text)
            except Exception as e:
                print("Error {} : ".format(e) )
                text = "识别失败,请重试"
            return render (request,'appsolve/spechtotext.html',{'l5arij':request.FILES['theFileINEEd'],'link':file_url,'fulllink':full_link, 'text': text})   
        else:
            return render (request,'appsolve/spechtotext.html',{'problem':"you must upload a file mp3"}) 
    else:
        return render (request,'appsolve/spechtotext.html',)

内容的提问来源于stack exchange,提问作者Amine Riyahi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.06 22:45:03