Gradio语音转录界面本地启动及运行异常问题求助
Gradio语音转录界面本地运行异常问题
问题概述
代码在Google Colab中运行正常,但本地命令行和PyCharm环境下表现异常,具体问题如下:
环境与依赖安装
已执行以下命令安装依赖:
pip install git+https://github.com/openai/whisper.gitpip install ffmpeg-python
命令行运行异常表现
执行脚本后程序直接挂起,终止进程后才显示本该正常输出的启动日志:
$ python test_audio_delay.py <--- 执行后挂起,直到进程被终止。 * Running on local URL: http://127.0.0.1:8000 * Running on public URL: https://097932757b4f98b0b0.gradio.live This share link expires in 72 hours. For free permanent hosting and GPU upgrades, run `gradio deploy` from the terminal in the working directory to deploy to Hugging Face Spaces (https://huggingface.co/spaces) Keyboard interruption in main thread... closing server. Killing tunnel 127.0.0.1:8000 <> https://097932757b4f98b0b0.gradio.live (.venv)
PyCharm运行异常表现
能捕获到语音录制生成的文件名,但访问本地URL处理音频时出现文件找不到错误:
C:\Git\Git\Documents\Workspace\RAG01\.venv\Scripts\python.exe C:\Git\Git\Documents\Workspace\RAG01\test_audio_delay.py * Running on local URL: http://127.0.0.1:8000 * Running on public URL: https://820b9f762f181a269b.gradio.live This share link expires in 72 hours. For free permanent hosting and GPU upgrades, run `gradio deploy` from the terminal in the working directory to deploy to Hugging Face Spaces (https://huggingface.co/spaces) Processing file: C:\Users\NaderAfshar\AppData\Local\Temp\gradio\4418c42db7d360241e0cde3c9dd706ef3ac72ad0fd5513f3bebb94f5936e4e75\audio.wav ... ... File "C:\Program Files\Python311\Lib\subprocess.py", line 1538, in _execute_child hp, ht, pid, tid = _winapi.CreateProcess(executable, args, ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ FileNotFoundError: [WinError 2] The system cannot find the file specified
代码示例
import gradio as gr import whisper import os def transcribe_audio(audio_file): if not os.path.exists(audio_file): print(f"Cannot locate file: {audio_file}") return "Error: Audio file not found!" else: print(f"Processing file: {audio_file}") model = whisper.load_model("base") result = model.transcribe(audio_file, fp16=False) return result["text"] def main(): audio_input = gr.Audio(sources=["microphone"], type="filepath"), output_text = gr.Textbox(label="Transcription") iface = gr.Interface(fn=transcribe_audio, inputs=audio_input, outputs=output_text, title="Audio Transcription App" ) iface.launch( share=True, debug=True, server_port=8000, prevent_thread_lock=True ) if __name__ == '__main__': main()
解决方案
1. 命令行挂起问题
问题根源是prevent_thread_lock=True参数,该参数会让Gradio在后台启动服务,主线程直接结束,导致命令行呈现挂起状态。解决方式二选一:
- 移除
prevent_thread_lock=True参数,Gradio会默认阻塞主线程,启动日志会实时输出,服务保持运行。 - 保留参数但添加主线程阻塞逻辑,在
iface.launch()后追加代码:import time while True: time.sleep(1)
2. PyCharm中FileNotFoundError问题
错误并非找不到音频文件,而是Whisper依赖的ffmpeg可执行文件未安装或未加入系统PATH。ffmpeg-python仅为Python绑定,不包含实际的ffmpeg程序:
- 下载对应系统的ffmpeg二进制包,解压后将包含
ffmpeg.exe的目录添加到系统环境变量PATH中。 - 重启PyCharm和终端,确保PATH配置生效。
- 验证:在命令行输入
ffmpeg -version,能正常输出版本信息即为配置成功。
3. 代码语法修正
audio_input定义末尾多了一个逗号,导致它被解析为元组,Gradio无法正确识别输入组件,需移除逗号:
audio_input = gr.Audio(sources=["microphone"], type="filepath") # 移除末尾逗号
内容的提问来源于stack exchange,提问作者Nader Afshar
相关产品推荐
相关产品推荐

