You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

OpenAI Whisper模块无法找到已存在音频文件问题求助

问题描述

在Windows 10系统中用Python开发基于OpenAI Whisper的音频转文本程序,脚本能正常获取同目录下audio.mp4的时间戳(证明文件存在),但传入绝对路径调用Whisper的transcribe方法时,抛出FileNotFoundError报错。

代码

import whisper
import pandas as pd
import os
import sys
from datetime import datetime

# Show the current working directory
cwd = os.getcwd()

print ("Current working directory: {0}\n".format(cwd))

# Transcript a previously downloaded audio file.
# audio_file = "./audio.mp4"

# with open(os.path.join(sys.path[0], "audio.mp4"), "r") as f:
audio_file = os.path.join(cwd, "audio.mp4")

print ("Using audio input file: {0}\n".format(audio_file))

# Get the timestamp for the file
timestamp = os.path.getmtime(audio_file)

# Convert the timestamp to a datetime object
dt = datetime.fromtimestamp(timestamp)

# Format the datetime object in the desired format
formatted_timestamp = dt.strftime("%m/%d/%Y")

# Print the formatted timestamp
print("Input file timestamp: {0}\n\n".format(formatted_timestamp))

#Load the OpenAI Whisper model
whisper_model = whisper.load_model("tiny")

# Transcribe the audio.
transcription = whisper_model.transcribe(audio_file)

# Display the transcription.  This will display 
#  the transcription result in segments with 
#  start and end time. The full concatenated 
#  string is available as transcription['text']

# print as DataFrame
df = pd.DataFrame(transcription['segments'], columns=['start', 'end', 'text'])
print(df)

# or, print as String
print(transcription['text'])

程序输出

C:\Users\main\Documents\GitHub\ME\open-ai\whisper\python-utilities>python transcribe-audio.py
Current working directory: C:\Users\main\Documents\GitHub\ME\open-ai\whisper\python-utilities

Using audio input file: C:\Users\main\Documents\GitHub\ME\open-ai\whisper\python-utilities\audio.mp4

Input file timestamp: 12/27/2022


C:\Python310\lib\site-packages\whisper\transcribe.py:78: UserWarning: FP16 is not supported on CPU; using FP32 instead
  warnings.warn("FP16 is not supported on CPU; using FP32 instead")
Traceback (most recent call last):
  File "C:\Users\main\Documents\GitHub\ME\open-ai\whisper\python-utilities\transcribe-audio.py", line 36, in <module>
    transcription = whisper_model.transcribe(audio_file)
  File "C:\Python310\lib\site-packages\whisper\transcribe.py", line 84, in transcribe
    mel = log_mel_spectrogram(audio)
  File "C:\Python310\lib\site-packages\whisper\audio.py", line 111, in log_mel_spectrogram
    audio = load_audio(audio)
  File "C:\Python310\lib\site-packages\whisper\audio.py", line 42, in load_audio
    ffmpeg.input(file, threads=0)
  File "C:\Python310\lib\site-packages\ffmpeg\_run.py", line 313, in run
    process = run_async(
  File "C:\Python310\lib\site-packages\ffmpeg\_run.py", line 284, in run_async
    return subprocess.Popen(
  File "C:\Python310\lib\subprocess.py", line 966, in __init__
    self._execute_child(args, executable, preexec_fn, close_fds,
  File "C:\Python310\lib\subprocess.py", line 1435, in _execute_child
    hp, ht, pid, tid = _winapi.CreateProcess(executable, args,
FileNotFoundError: [WinError 2] The system cannot find the file specified

问题分析

报错并非找不到audio.mp4文件(脚本已成功读取文件时间戳),而是Whisper依赖的ffmpeg工具未被系统找到。从报错栈可见,错误发生在调用ffmpeg.input()的subprocess环节,Windows环境下若ffmpeg未安装或未添加到系统PATH环境变量,Python的subprocess模块无法定位ffmpeg可执行文件,从而抛出FileNotFoundError。

修复方法

方法1:安装并配置ffmpeg到系统PATH

  1. 下载ffmpeg:获取Windows版本的ffmpeg压缩包,解压到任意目录(如C:\ffmpeg)。
  2. 添加环境变量:打开系统环境变量设置,将ffmpeg解压目录下的bin文件夹路径(如C:\ffmpeg\bin)添加到系统变量的Path中。
  3. 验证配置:打开新的命令提示符窗口,输入ffmpeg -version,若能正常显示ffmpeg版本信息,说明配置成功。
  4. 重启Python运行环境(命令提示符、IDE等),重新运行脚本。

方法2:在代码中指定ffmpeg路径

若不想修改系统PATH,可在代码中手动指定ffmpeg可执行文件的路径:

# 导入whisper前设置ffmpeg路径
import os
os.environ["FFMPEG_BINARY"] = "C:/ffmpeg/bin/ffmpeg.exe"

import whisper
# 后续代码保持不变

注意:路径需使用正斜杠或双反斜杠作为分隔符,确保指向正确的ffmpeg可执行文件。


内容的提问来源于stack exchange,提问作者Robert Oschler

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.06 15:20:29