You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Gradio语音转录界面本地启动及运行异常问题求助

Gradio语音转录界面本地运行异常问题

问题概述

代码在Google Colab中运行正常,但本地命令行和PyCharm环境下表现异常,具体问题如下:

环境与依赖安装

已执行以下命令安装依赖:

  • pip install git+https://github.com/openai/whisper.git
  • pip install ffmpeg-python

命令行运行异常表现

执行脚本后程序直接挂起,终止进程后才显示本该正常输出的启动日志:

$ python test_audio_delay.py  <--- 执行后挂起,直到进程被终止。
* Running on local URL:  http://127.0.0.1:8000
* Running on public URL: https://097932757b4f98b0b0.gradio.live

This share link expires in 72 hours. For free permanent hosting and GPU upgrades, run `gradio deploy` from the terminal in the working directory to deploy to Hugging Face Spaces (https://huggingface.co/spaces)
Keyboard interruption in main thread... closing server.
Killing tunnel 127.0.0.1:8000 <> https://097932757b4f98b0b0.gradio.live
(.venv)

PyCharm运行异常表现

能捕获到语音录制生成的文件名,但访问本地URL处理音频时出现文件找不到错误:

C:\Git\Git\Documents\Workspace\RAG01\.venv\Scripts\python.exe C:\Git\Git\Documents\Workspace\RAG01\test_audio_delay.py 
* Running on local URL:  http://127.0.0.1:8000
* Running on public URL: https://820b9f762f181a269b.gradio.live

This share link expires in 72 hours. For free permanent hosting and GPU upgrades, run `gradio deploy` from the terminal in the working directory to deploy to Hugging Face Spaces (https://huggingface.co/spaces)

Processing file: C:\Users\NaderAfshar\AppData\Local\Temp\gradio\4418c42db7d360241e0cde3c9dd706ef3ac72ad0fd5513f3bebb94f5936e4e75\audio.wav
...
...
File "C:\Program Files\Python311\Lib\subprocess.py", line 1538, in _execute_child
    hp, ht, pid, tid = _winapi.CreateProcess(executable, args,
                       ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
FileNotFoundError: [WinError 2] The system cannot find the file specified

代码示例

import gradio as gr
import whisper
import os


def transcribe_audio(audio_file):

    if not os.path.exists(audio_file):
        print(f"Cannot locate file: {audio_file}")
        return "Error: Audio file not found!"
    else:
        print(f"Processing file: {audio_file}")

    model = whisper.load_model("base")
    result = model.transcribe(audio_file, fp16=False)
    return result["text"]


def main():
    audio_input = gr.Audio(sources=["microphone"], type="filepath"),
    output_text = gr.Textbox(label="Transcription")

    iface = gr.Interface(fn=transcribe_audio,
                         inputs=audio_input,
                         outputs=output_text,
                         title="Audio Transcription App"
                         )

    iface.launch(
        share=True,
        debug=True,
        server_port=8000,
        prevent_thread_lock=True
    )


if __name__ == '__main__':
    main()

解决方案

1. 命令行挂起问题

问题根源是prevent_thread_lock=True参数,该参数会让Gradio在后台启动服务,主线程直接结束,导致命令行呈现挂起状态。解决方式二选一:

  • 移除prevent_thread_lock=True参数,Gradio会默认阻塞主线程,启动日志会实时输出,服务保持运行。
  • 保留参数但添加主线程阻塞逻辑,在iface.launch()后追加代码:
    import time
    while True:
        time.sleep(1)
    

2. PyCharm中FileNotFoundError问题

错误并非找不到音频文件,而是Whisper依赖的ffmpeg可执行文件未安装或未加入系统PATH。ffmpeg-python仅为Python绑定,不包含实际的ffmpeg程序:

  • 下载对应系统的ffmpeg二进制包,解压后将包含ffmpeg.exe的目录添加到系统环境变量PATH中。
  • 重启PyCharm和终端,确保PATH配置生效。
  • 验证:在命令行输入ffmpeg -version,能正常输出版本信息即为配置成功。

3. 代码语法修正

audio_input定义末尾多了一个逗号,导致它被解析为元组,Gradio无法正确识别输入组件,需移除逗号:

audio_input = gr.Audio(sources=["microphone"], type="filepath")  # 移除末尾逗号

内容的提问来源于stack exchange,提问作者Nader Afshar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.13 14:39:51