如何在Google Colab中用Hugging Face MMS模型实现语音转文本?
在Google Colab部署Facebook MMS语音转文本模型的完整步骤
一、已完成的依赖安装(可跳过)
你已经安装好所需依赖,对应命令如下:
!pip install transformers !pip install datasets>=2.6.1 !pip install git+https://github.com/huggingface/transformers !pip install librosa !pip install evaluate>=0.30 !pip install jiwer !pip install gradio
二、核心实施步骤
1. 加载MMS预训练模型与处理器
根据目标语言选择对应模型(比如中文选facebook/mms-1b-zh,通用多语言选facebook/mms-1b-all),加载模型和语音处理器:
from transformers import AutoModelForSpeechSeq2Seq, AutoProcessor, pipeline # 替换为你需要的模型名称 model_name = "facebook/mms-1b-all" processor = AutoProcessor.from_pretrained(model_name) model = AutoModelForSpeechSeq2Seq.from_pretrained(model_name, device_map="auto") # 构建语音转文本管道 asr_pipeline = pipeline( "automatic-speech-recognition", model=model, tokenizer=processor.tokenizer, feature_extractor=processor.feature_extractor, device_map="auto" )
2. 加载并预处理自定义语音数据
将语音文件上传到Colab(可通过左侧文件面板上传,或挂载Google Drive读取),然后用librosa统一音频采样率为模型要求的16kHz:
import librosa # 加载单音频文件,采样率转为16kHz audio_path = "你的语音文件路径.wav" audio, sr = librosa.load(audio_path, sr=16000) # 批量加载文件夹内的音频文件 import os audio_folder = "你的语音文件夹路径" audio_files = [os.path.join(audio_folder, f) for f in os.listdir(audio_folder) if f.endswith(".wav")]
3. 执行语音转文本推理
单文件或批量执行推理,输出转写结果:
# 单文件推理 result = asr_pipeline(audio) print("转写结果:", result["text"]) # 批量推理 for file in audio_files: audio, sr = librosa.load(file, sr=16000) result = asr_pipeline(audio) print(f"{file} 转写结果:", result["text"])
4. (可选)评估模型转写精度
用jiwer计算词错误率(WER),评估模型在你的数据集上的性能(需提前准备对应语音的标注文本):
from evaluate import load wer_metric = load("wer") # 假设你有标注文本列表labels,和转写结果列表predictions labels = ["标注文本1", "标注文本2"] predictions = [result1["text"], result2["text"]] wer_score = wer_metric.compute(predictions=predictions, references=labels) print(f"词错误率(WER):{wer_score}")
5. (可选)搭建可视化Demo
用Gradio快速搭建可交互的语音转文本界面,方便测试:
import gradio as gr import librosa def transcribe_audio(audio): sr, audio_data = audio # 转为16kHz单声道 audio_data = librosa.resample(audio_data, orig_sr=sr, target_sr=16000) result = asr_pipeline(audio_data) return result["text"] demo = gr.Interface( fn=transcribe_audio, inputs=gr.Audio(type="numpy"), outputs=gr.Textbox(label="转写结果"), title="MMS语音转文本Demo" ) demo.launch(share=True)
三、关键优化建议
- 选对语言模型:优先选择目标语言专属的MMS模型(比如
facebook/mms-1b-zh对应中文),比通用多语言模型转写精度更高 - 显存优化:Colab免费版显存有限,加载模型时可添加
load_in_8bit=True启用8位量化,减少显存占用 - 音频预处理:确保所有输入音频统一为16kHz采样率、单声道,否则会影响转写效果
- 批量处理优化:大规模数据集处理时,可分批次推理,避免内存溢出
内容的提问来源于stack exchange,提问作者coder
相关产品推荐
相关产品推荐

