You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Google Colab中用Hugging Face MMS模型实现语音转文本?

在Google Colab部署Facebook MMS语音转文本模型的完整步骤

一、已完成的依赖安装(可跳过)

你已经安装好所需依赖,对应命令如下:

!pip install transformers
!pip install datasets>=2.6.1
!pip install git+https://github.com/huggingface/transformers
!pip install librosa
!pip install evaluate>=0.30
!pip install jiwer
!pip install gradio

二、核心实施步骤

1. 加载MMS预训练模型与处理器

根据目标语言选择对应模型(比如中文选facebook/mms-1b-zh,通用多语言选facebook/mms-1b-all),加载模型和语音处理器:

from transformers import AutoModelForSpeechSeq2Seq, AutoProcessor, pipeline

# 替换为你需要的模型名称
model_name = "facebook/mms-1b-all"
processor = AutoProcessor.from_pretrained(model_name)
model = AutoModelForSpeechSeq2Seq.from_pretrained(model_name, device_map="auto")

# 构建语音转文本管道
asr_pipeline = pipeline(
    "automatic-speech-recognition",
    model=model,
    tokenizer=processor.tokenizer,
    feature_extractor=processor.feature_extractor,
    device_map="auto"
)

2. 加载并预处理自定义语音数据

将语音文件上传到Colab(可通过左侧文件面板上传,或挂载Google Drive读取),然后用librosa统一音频采样率为模型要求的16kHz:

import librosa

# 加载单音频文件,采样率转为16kHz
audio_path = "你的语音文件路径.wav"
audio, sr = librosa.load(audio_path, sr=16000)

# 批量加载文件夹内的音频文件
import os
audio_folder = "你的语音文件夹路径"
audio_files = [os.path.join(audio_folder, f) for f in os.listdir(audio_folder) if f.endswith(".wav")]

3. 执行语音转文本推理

单文件或批量执行推理,输出转写结果:

# 单文件推理
result = asr_pipeline(audio)
print("转写结果:", result["text"])

# 批量推理
for file in audio_files:
    audio, sr = librosa.load(file, sr=16000)
    result = asr_pipeline(audio)
    print(f"{file} 转写结果:", result["text"])

4. (可选)评估模型转写精度

用jiwer计算词错误率(WER),评估模型在你的数据集上的性能(需提前准备对应语音的标注文本):

from evaluate import load

wer_metric = load("wer")

# 假设你有标注文本列表labels,和转写结果列表predictions
labels = ["标注文本1", "标注文本2"]
predictions = [result1["text"], result2["text"]]

wer_score = wer_metric.compute(predictions=predictions, references=labels)
print(f"词错误率(WER):{wer_score}")

5. (可选)搭建可视化Demo

用Gradio快速搭建可交互的语音转文本界面,方便测试:

import gradio as gr
import librosa

def transcribe_audio(audio):
    sr, audio_data = audio
    # 转为16kHz单声道
    audio_data = librosa.resample(audio_data, orig_sr=sr, target_sr=16000)
    result = asr_pipeline(audio_data)
    return result["text"]

demo = gr.Interface(
    fn=transcribe_audio,
    inputs=gr.Audio(type="numpy"),
    outputs=gr.Textbox(label="转写结果"),
    title="MMS语音转文本Demo"
)

demo.launch(share=True)

三、关键优化建议

  • 选对语言模型:优先选择目标语言专属的MMS模型(比如facebook/mms-1b-zh对应中文),比通用多语言模型转写精度更高
  • 显存优化:Colab免费版显存有限,加载模型时可添加load_in_8bit=True启用8位量化,减少显存占用
  • 音频预处理:确保所有输入音频统一为16kHz采样率、单声道,否则会影响转写效果
  • 批量处理优化:大规模数据集处理时,可分批次推理,避免内存溢出

内容的提问来源于stack exchange,提问作者coder

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.15 05:09:55