You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python新手使用whisper-timestamped处理MP4遇WinError 2报错求助

问题:whisper-timestamped加载MP4文件报错[WinError 2]
  • 本人Python新手,尝试用whisper-timestamped库从MP4文件生成字幕,但执行audio = whisper_timestamped.load_audio(videofile)时一直报错[WinError 2] The system cannot find the file specified
  • 奇怪的是,用moviepy打开同一个MP4文件并转换保存音频的代码可以正常运行
  • 完整报错信息:
FileNotFoundError                         Traceback (most recent call last)
Cell In[6], line 1
----> 1 audio = whisper_timestamped.load_audio("bausenvid.mp4")
      2 model = whisper_timestamped.load_model('base',device='gpu')
      3 results = whisper_timestamped.transcribe(model,audio,language='en')

File ~\AppData\Local\Programs\Python\Python310\lib\site-packages\whisper\audio.py:58, in load_audio(file, sr)
     56 # fmt: on
     57 try:
---> 58     out = run(cmd, capture_output=True, check=True).stdout
     59 except CalledProcessError as e:
     60     raise RuntimeError(f"Failed to load audio: {e.stderr.decode()}") from e

File ~\AppData\Local\Programs\Python\Python310\lib\subprocess.py:501, in run(input, capture_output, timeout, check, *popenargs, **kwargs)
    498     kwargs['stdout'] = PIPE
    499     kwargs['stderr'] = PIPE
---> 501 with Popen(*popenargs, **kwargs) as process:
    502     try:
    503         stdout, stderr = process.communicate(input, timeout=timeout)

File ~\AppData\Local\Programs\Python\Python310\lib\subprocess.py:966, in Popen.__init__(self, args, bufsize, executable, stdin, stdout, stderr, preexec_fn, close_fds, shell, cwd, env, universal_newlines, startupinfo, creationflags, restore_signals, start_new_session, pass_fds, user, group, extra_groups, encoding, errors, text, umask, pipesize)
    962         if self.text_mode:
    963             self.stderr = io.TextIOWrapper(self.stderr,
    964                     encoding=encoding, errors=errors)
---> 966     self._execute_child(args, executable, preexec_fn, close_fds,
    967                         pass_fds, cwd, env,
    968                         startupinfo, creationflags, shell,
    969                         p2cread, p2cwrite,
    970                         c2pread, c2pwrite,
    971                         errread, errwrite,
    972                         restore_signals,
    973                         gid, gids, uid, umask,
    974                         start_new_session)
    975 except:
    976     # Cleanup if the child failed starting.
    977     for f in filter(None, (self.stdin, self.stdout, self.stderr)):

File ~\AppData\Local\Programs\Python\Python310\lib\subprocess.py:1435, in Popen._execute_child(self, args, executable, preexec_fn, close_fds, pass_fds, cwd, env, startupinfo, creationflags, shell, p2cread, p2cwrite, c2pread, c2pwrite, errread, errwrite, unused_restore_signals, unused_gid, unused_gids, unused_uid, unused_umask, unused_start_new_session)
   1433 # Start the process
   1434 try:
---> 1435     hp, ht, pid, tid = _winapi.CreateProcess(executable, args,
   1436                              # no special security
   1437                              None, None,
   1438                              int(not close_fds),
   1439                              creationflags,
   1440                              env,
   1441                              cwd,
   1442                              startupinfo)
   1443 finally:
   1444     # Child is launched. Close the parent's copy of those pipe
   1445     # handles that only the child should have open.  You need
   (...)
   1448     # pipe will not close when the child process exits and the
   1449     # ReadFile will hang.
   1450     self._close_pipe_fds(p2cread, p2cwrite,
   1451                          c2pread, c2pwrite,
   1452                          errread, errwrite)

FileNotFoundError: [WinError 2] The system cannot find the file specified

解决方法

核心原因

whisper-timestamped的load_audio底层依赖ffmpeg工具处理音视频文件,这个报错不是找不到你的MP4文件,而是系统找不到ffmpeg的执行程序。moviepy要么自带了ffmpeg,要么已经正确配置了ffmpeg路径,所以能正常运行。

方案1:安装并配置ffmpeg

  1. Windows用户安装方式:
    • 手动下载:下载ffmpeg静态编译包,解压后找到bin目录(内含ffmpeg.exe),将该目录添加到系统环境变量PATH中
    • 包管理器安装:用Chocolatey执行choco install ffmpeg,或Scoop执行scoop install ffmpeg(需先安装对应包管理器)
  2. 验证配置:打开新的命令提示符,输入ffmpeg -version,能显示版本信息即为配置成功
  3. 重启Python环境(如Jupyter Notebook、VSCode等)后重新运行代码

方案2:用moviepy先提取音频再传给whisper-timestamped

如果不想折腾ffmpeg配置,直接用你已能正常运行的moviepy先提取音频,再转换为whisper兼容的格式:

from moviepy.editor import VideoFileClip
import whisper_timestamped
import numpy as np

# 用moviepy读取视频并提取音频
video_clip = VideoFileClip("bausenvid.mp4")
# 转换为whisper默认的16000Hz采样率并转为单声道
audio_array = video_clip.audio.to_soundarray(fps=16000)
audio_mono = np.mean(audio_array, axis=1)

# 正常调用whisper-timestamped
model = whisper_timestamped.load_model('base', device='gpu')
results = whisper_timestamped.transcribe(model, audio_mono, language='en')

内容的提问来源于stack exchange,提问作者Abdo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.28 14:24:59