Django语音转文本视图报错:音频无法被读取为PCM WAV等支持格式
问题复现
我的语音转文本视图无法正常运行,已有wav格式文件,但运行时报出如下错误:
Traceback (most recent call last): File "C:\Users\privet01\Desktop\python projects\projex x\myvev\lib\site-packages\django\core\handlers\exception.py", line 47, in inner response = get_response(request) File "C:\Users\privet01\Desktop\python projects\projex x\myvev\lib\site-packages\django\core\handlers\base.py", line 179, in _get_response response = wrapped_callback(request, *callback_args, **callback_kwargs) File "C:\Users\privet01\Desktop\python projects\projex x\alpha1\appsolve\views.py", line 636, in spechtotext with sr.AudioFile(sound) as source: File "C:\Users\privet01\Desktop\python projects\projex x\myvev\lib\site-packages\speech_recognition\__init__.py", line 236, in __enter__ raise ValueError("Audio file could not be read as PCM WAV, AIFF/AIFF-C, or Native FLAC; check if file is corrupted or in another format") ValueError: Audio file could not be read as PCM WAV, AIFF/AIFF-C, or Native FLAC; check if file is corrupted or in another format [23/Aug/2021 13:26:44] "POST /spechtotext/ HTTP/1.1" 500 152359
原视图代码
import speech_recognition as sr from django.core.files.storage import FileSystemStorage def spechtotext(request) : if request.method=='POST': if request.FILES['theFileINEEd'] and '.mp3' in str(request.FILES['theFileINEEd'].name).lower() : myfile = request.FILES['theFileINEEd'] fs = FileSystemStorage('') filename = fs.save(str(request.user)+"speachtotext.WAV", myfile) file_url = fs.url(filename) full_link='http://127.0.0.1:8000'+file_url # changeble to life mode r = sr.Recognizer() print(type(filename)) print('hhhhhhhhhhhhhhhhhhhhhhhhhhhhhhhhhhhhhhhhhhhhhhhhhhhhhhhhhh') r = sr.Recognizer() import sys print(sys.version_info) print('qqqqqqqqqqqqqqqqqqqqqqqqq'+str(filename)+'aaaaaaaaaaaaaaaaaaaa') hellow=sr.AudioFile(str(filename)) sound = str(filename) with sr.AudioFile(sound) as source: r.adjust_for_ambient_noise(source) print("Converting Audio To Text ..... ") audio = r.listen(source) try: print("Converted Audio Is : \n" + r.recognize_google(audio)) except Exception as e: print("Error {} : ".format(e) ) audio_text = r.listen(source) try: # using google speech recognition text = r.recognize_google(audio_text) print('Converting audio transcripts into text ...') print("this is the text : ",text) except: print('Sorry.. run again...') return render (request,'appsolve/spechtotext.html',{'l5arij':request.FILES['theFileINEEd'],'link':file_url,'fulllink':full_link}) else: return render (request,'appsolve/spechtotext.html',{'problem':"you must upload a file mp3"}) else: return render (request,'appsolve/spechtotext.html',)
相关材料
- 网页报错截图:

- 文件路径截图:

已尝试操作
- 确认音频文件已存储在对应路径下
- 分别测试了mp3和PCM格式的音频
- 使用音质清晰的音频测试,问题仍未解决
解决方案
根因说明
你代码中直接将上传的mp3文件改后缀保存为.WAV格式,仅修改后缀不会改变文件本身的编码结构,speech_recognition库的AudioFile仅支持原生PCM编码的WAV、AIFF、FLAC格式,无法识别改了后缀的mp3文件,因此抛出对应错误。另外你代码中except块里调用r.listen(source)时,source已经被with块关闭,会额外触发报错。
修复步骤
- 安装依赖库
你需要用pydub完成音频转码,先安装相关依赖:
pip install pydub
windows环境需要额外下载ffmpeg工具,将其bin目录添加到系统环境变量Path中,或在代码中指定ffmpeg路径。
- 修改视图代码
核心修改音频存储和转码逻辑,修复异常块的错误逻辑:
import speech_recognition as sr from django.core.files.storage import FileSystemStorage from pydub import AudioSegment import os def spechtotext(request): if request.method=='POST': if request.FILES['theFileINEEd'] and '.mp3' in str(request.FILES['theFileINEEd'].name).lower() : myfile = request.FILES['theFileINEEd'] fs = FileSystemStorage('') # 先保存临时mp3文件 temp_mp3_name = f"{str(request.user)}_temp.mp3" temp_mp3_path = fs.save(temp_mp3_name, myfile) # 转码为PCM编码的WAV文件,转成单声道16k采样率适配语音识别 audio = AudioSegment.from_mp3(temp_mp3_path) wav_name = f"{str(request.user)}_speachtotext.WAV" audio.export(wav_name, format="wav", parameters=["-ac", "1", "-ar", "16000"]) # 删除临时mp3文件 os.remove(temp_mp3_path) file_url = fs.url(wav_name) full_link='http://127.0.0.1:8000'+file_url r = sr.Recognizer() with sr.AudioFile(wav_name) as source: r.adjust_for_ambient_noise(source) print("Converting Audio To Text ..... ") audio = r.listen(source) try: text = r.recognize_google(audio, language="zh-CN") # 可指定识别语言,默认是英文 print("Converted Audio Is : \n" + text) except Exception as e: print("Error {} : ".format(e) ) text = "识别失败,请重试" return render (request,'appsolve/spechtotext.html',{'l5arij':request.FILES['theFileINEEd'],'link':file_url,'fulllink':full_link, 'text': text}) else: return render (request,'appsolve/spechtotext.html',{'problem':"you must upload a file mp3"}) else: return render (request,'appsolve/spechtotext.html',)
内容的提问来源于stack exchange,提问作者Amine Riyahi
相关产品推荐
相关产品推荐

