如何在Google Colab中使用Hugging Face微调Whisper模型实现印度语言音频转写
在Google Colab中使用thennal/whisper-medium-ml转写印度语言音频
步骤1:配置Colab运行环境
- 打开Google Colab,新建一个笔记本
- 点击顶部菜单栏「修改」→「笔记本设置」,在「硬件加速器」下拉菜单选择「GPU」,点击「保存」
步骤2:安装依赖库
在Colab的代码单元格中运行以下命令:
!pip install --upgrade pip !pip install transformers datasets torch soundfile
步骤3:加载模型与处理器
运行以下代码加载目标模型及音频处理器:
from transformers import AutoModelForSpeechSeq2Seq, AutoProcessor, pipeline import torch model_name = "thennal/whisper-medium-ml" processor = AutoProcessor.from_pretrained(model_name) model = AutoModelForSpeechSeq2Seq.from_pretrained(model_name) # 自动使用GPU(如果可用) device = "cuda:0" if torch.cuda.is_available() else "cpu" model.to(device) # 创建转写管道 pipe = pipeline( "automatic-speech-recognition", model=model, tokenizer=processor.tokenizer, feature_extractor=processor.feature_extractor, device=device, )
步骤4:上传本地音频文件
运行以下代码,在弹出窗口中选择你的印度语言音频文件(支持MP3、WAV等格式):
from google.colab import files uploaded = files.upload() audio_file = next(iter(uploaded.keys()))
步骤5:执行音频转写
运行以下代码获取转写结果:
result = pipe(audio_file) print("转写文本:") print(result["text"])
注意事项
- 长音频转写需要等待几分钟,请耐心等候
- 音频清晰度越高,转写准确率越好,尽量避免背景噪音过大的文件
内容的提问来源于stack exchange,提问作者user22347502
相关产品推荐
相关产品推荐

