You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Google Colab中使用Hugging Face微调Whisper模型实现印度语言音频转写

在Google Colab中使用thennal/whisper-medium-ml转写印度语言音频

步骤1:配置Colab运行环境

  • 打开Google Colab,新建一个笔记本
  • 点击顶部菜单栏「修改」→「笔记本设置」,在「硬件加速器」下拉菜单选择「GPU」,点击「保存」

步骤2:安装依赖库

在Colab的代码单元格中运行以下命令:

!pip install --upgrade pip
!pip install transformers datasets torch soundfile

步骤3:加载模型与处理器

运行以下代码加载目标模型及音频处理器:

from transformers import AutoModelForSpeechSeq2Seq, AutoProcessor, pipeline
import torch

model_name = "thennal/whisper-medium-ml"
processor = AutoProcessor.from_pretrained(model_name)
model = AutoModelForSpeechSeq2Seq.from_pretrained(model_name)

# 自动使用GPU(如果可用)
device = "cuda:0" if torch.cuda.is_available() else "cpu"
model.to(device)

# 创建转写管道
pipe = pipeline(
    "automatic-speech-recognition",
    model=model,
    tokenizer=processor.tokenizer,
    feature_extractor=processor.feature_extractor,
    device=device,
)

步骤4:上传本地音频文件

运行以下代码,在弹出窗口中选择你的印度语言音频文件(支持MP3、WAV等格式):

from google.colab import files

uploaded = files.upload()
audio_file = next(iter(uploaded.keys()))

步骤5:执行音频转写

运行以下代码获取转写结果:

result = pipe(audio_file)

print("转写文本:")
print(result["text"])

注意事项

  • 长音频转写需要等待几分钟,请耐心等候
  • 音频清晰度越高,转写准确率越好,尽量避免背景噪音过大的文件

内容的提问来源于stack exchange,提问作者user22347502

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.14 06:14:57