Python新手使用whisper-timestamped处理MP4遇WinError 2报错求助
问题:whisper-timestamped加载MP4文件报错[WinError 2]
- 本人Python新手,尝试用whisper-timestamped库从MP4文件生成字幕,但执行
audio = whisper_timestamped.load_audio(videofile)时一直报错[WinError 2] The system cannot find the file specified - 奇怪的是,用moviepy打开同一个MP4文件并转换保存音频的代码可以正常运行
- 完整报错信息:
FileNotFoundError Traceback (most recent call last) Cell In[6], line 1 ----> 1 audio = whisper_timestamped.load_audio("bausenvid.mp4") 2 model = whisper_timestamped.load_model('base',device='gpu') 3 results = whisper_timestamped.transcribe(model,audio,language='en') File ~\AppData\Local\Programs\Python\Python310\lib\site-packages\whisper\audio.py:58, in load_audio(file, sr) 56 # fmt: on 57 try: ---> 58 out = run(cmd, capture_output=True, check=True).stdout 59 except CalledProcessError as e: 60 raise RuntimeError(f"Failed to load audio: {e.stderr.decode()}") from e File ~\AppData\Local\Programs\Python\Python310\lib\subprocess.py:501, in run(input, capture_output, timeout, check, *popenargs, **kwargs) 498 kwargs['stdout'] = PIPE 499 kwargs['stderr'] = PIPE ---> 501 with Popen(*popenargs, **kwargs) as process: 502 try: 503 stdout, stderr = process.communicate(input, timeout=timeout) File ~\AppData\Local\Programs\Python\Python310\lib\subprocess.py:966, in Popen.__init__(self, args, bufsize, executable, stdin, stdout, stderr, preexec_fn, close_fds, shell, cwd, env, universal_newlines, startupinfo, creationflags, restore_signals, start_new_session, pass_fds, user, group, extra_groups, encoding, errors, text, umask, pipesize) 962 if self.text_mode: 963 self.stderr = io.TextIOWrapper(self.stderr, 964 encoding=encoding, errors=errors) ---> 966 self._execute_child(args, executable, preexec_fn, close_fds, 967 pass_fds, cwd, env, 968 startupinfo, creationflags, shell, 969 p2cread, p2cwrite, 970 c2pread, c2pwrite, 971 errread, errwrite, 972 restore_signals, 973 gid, gids, uid, umask, 974 start_new_session) 975 except: 976 # Cleanup if the child failed starting. 977 for f in filter(None, (self.stdin, self.stdout, self.stderr)): File ~\AppData\Local\Programs\Python\Python310\lib\subprocess.py:1435, in Popen._execute_child(self, args, executable, preexec_fn, close_fds, pass_fds, cwd, env, startupinfo, creationflags, shell, p2cread, p2cwrite, c2pread, c2pwrite, errread, errwrite, unused_restore_signals, unused_gid, unused_gids, unused_uid, unused_umask, unused_start_new_session) 1433 # Start the process 1434 try: ---> 1435 hp, ht, pid, tid = _winapi.CreateProcess(executable, args, 1436 # no special security 1437 None, None, 1438 int(not close_fds), 1439 creationflags, 1440 env, 1441 cwd, 1442 startupinfo) 1443 finally: 1444 # Child is launched. Close the parent's copy of those pipe 1445 # handles that only the child should have open. You need (...) 1448 # pipe will not close when the child process exits and the 1449 # ReadFile will hang. 1450 self._close_pipe_fds(p2cread, p2cwrite, 1451 c2pread, c2pwrite, 1452 errread, errwrite) FileNotFoundError: [WinError 2] The system cannot find the file specified
解决方法
核心原因
whisper-timestamped的load_audio底层依赖ffmpeg工具处理音视频文件,这个报错不是找不到你的MP4文件,而是系统找不到ffmpeg的执行程序。moviepy要么自带了ffmpeg,要么已经正确配置了ffmpeg路径,所以能正常运行。
方案1:安装并配置ffmpeg
- Windows用户安装方式:
- 手动下载:下载ffmpeg静态编译包,解压后找到
bin目录(内含ffmpeg.exe),将该目录添加到系统环境变量PATH中 - 包管理器安装:用Chocolatey执行
choco install ffmpeg,或Scoop执行scoop install ffmpeg(需先安装对应包管理器)
- 手动下载:下载ffmpeg静态编译包,解压后找到
- 验证配置:打开新的命令提示符,输入
ffmpeg -version,能显示版本信息即为配置成功 - 重启Python环境(如Jupyter Notebook、VSCode等)后重新运行代码
方案2:用moviepy先提取音频再传给whisper-timestamped
如果不想折腾ffmpeg配置,直接用你已能正常运行的moviepy先提取音频,再转换为whisper兼容的格式:
from moviepy.editor import VideoFileClip import whisper_timestamped import numpy as np # 用moviepy读取视频并提取音频 video_clip = VideoFileClip("bausenvid.mp4") # 转换为whisper默认的16000Hz采样率并转为单声道 audio_array = video_clip.audio.to_soundarray(fps=16000) audio_mono = np.mean(audio_array, axis=1) # 正常调用whisper-timestamped model = whisper_timestamped.load_model('base', device='gpu') results = whisper_timestamped.transcribe(model, audio_mono, language='en')
内容的提问来源于stack exchange,提问作者Abdo
相关产品推荐
相关产品推荐

