OpenAI Whisper模块无法找到已存在音频文件问题求助
问题描述
在Windows 10系统中用Python开发基于OpenAI Whisper的音频转文本程序,脚本能正常获取同目录下audio.mp4的时间戳(证明文件存在),但传入绝对路径调用Whisper的transcribe方法时,抛出FileNotFoundError报错。
代码
import whisper import pandas as pd import os import sys from datetime import datetime # Show the current working directory cwd = os.getcwd() print ("Current working directory: {0}\n".format(cwd)) # Transcript a previously downloaded audio file. # audio_file = "./audio.mp4" # with open(os.path.join(sys.path[0], "audio.mp4"), "r") as f: audio_file = os.path.join(cwd, "audio.mp4") print ("Using audio input file: {0}\n".format(audio_file)) # Get the timestamp for the file timestamp = os.path.getmtime(audio_file) # Convert the timestamp to a datetime object dt = datetime.fromtimestamp(timestamp) # Format the datetime object in the desired format formatted_timestamp = dt.strftime("%m/%d/%Y") # Print the formatted timestamp print("Input file timestamp: {0}\n\n".format(formatted_timestamp)) #Load the OpenAI Whisper model whisper_model = whisper.load_model("tiny") # Transcribe the audio. transcription = whisper_model.transcribe(audio_file) # Display the transcription. This will display # the transcription result in segments with # start and end time. The full concatenated # string is available as transcription['text'] # print as DataFrame df = pd.DataFrame(transcription['segments'], columns=['start', 'end', 'text']) print(df) # or, print as String print(transcription['text'])
程序输出
C:\Users\main\Documents\GitHub\ME\open-ai\whisper\python-utilities>python transcribe-audio.py Current working directory: C:\Users\main\Documents\GitHub\ME\open-ai\whisper\python-utilities Using audio input file: C:\Users\main\Documents\GitHub\ME\open-ai\whisper\python-utilities\audio.mp4 Input file timestamp: 12/27/2022 C:\Python310\lib\site-packages\whisper\transcribe.py:78: UserWarning: FP16 is not supported on CPU; using FP32 instead warnings.warn("FP16 is not supported on CPU; using FP32 instead") Traceback (most recent call last): File "C:\Users\main\Documents\GitHub\ME\open-ai\whisper\python-utilities\transcribe-audio.py", line 36, in <module> transcription = whisper_model.transcribe(audio_file) File "C:\Python310\lib\site-packages\whisper\transcribe.py", line 84, in transcribe mel = log_mel_spectrogram(audio) File "C:\Python310\lib\site-packages\whisper\audio.py", line 111, in log_mel_spectrogram audio = load_audio(audio) File "C:\Python310\lib\site-packages\whisper\audio.py", line 42, in load_audio ffmpeg.input(file, threads=0) File "C:\Python310\lib\site-packages\ffmpeg\_run.py", line 313, in run process = run_async( File "C:\Python310\lib\site-packages\ffmpeg\_run.py", line 284, in run_async return subprocess.Popen( File "C:\Python310\lib\subprocess.py", line 966, in __init__ self._execute_child(args, executable, preexec_fn, close_fds, File "C:\Python310\lib\subprocess.py", line 1435, in _execute_child hp, ht, pid, tid = _winapi.CreateProcess(executable, args, FileNotFoundError: [WinError 2] The system cannot find the file specified
问题分析
报错并非找不到audio.mp4文件(脚本已成功读取文件时间戳),而是Whisper依赖的ffmpeg工具未被系统找到。从报错栈可见,错误发生在调用ffmpeg.input()的subprocess环节,Windows环境下若ffmpeg未安装或未添加到系统PATH环境变量,Python的subprocess模块无法定位ffmpeg可执行文件,从而抛出FileNotFoundError。
修复方法
方法1:安装并配置ffmpeg到系统PATH
- 下载ffmpeg:获取Windows版本的ffmpeg压缩包,解压到任意目录(如
C:\ffmpeg)。 - 添加环境变量:打开系统环境变量设置,将ffmpeg解压目录下的
bin文件夹路径(如C:\ffmpeg\bin)添加到系统变量的Path中。 - 验证配置:打开新的命令提示符窗口,输入
ffmpeg -version,若能正常显示ffmpeg版本信息,说明配置成功。 - 重启Python运行环境(命令提示符、IDE等),重新运行脚本。
方法2:在代码中指定ffmpeg路径
若不想修改系统PATH,可在代码中手动指定ffmpeg可执行文件的路径:
# 导入whisper前设置ffmpeg路径 import os os.environ["FFMPEG_BINARY"] = "C:/ffmpeg/bin/ffmpeg.exe" import whisper # 后续代码保持不变
注意:路径需使用正斜杠或双反斜杠作为分隔符,确保指向正确的ffmpeg可执行文件。
内容的提问来源于stack exchange,提问作者Robert Oschler
相关产品推荐
相关产品推荐

