如何在OpenAI Whisper语音识别中获取单词级时间戳?
用OpenAI Whisper Python库获取单词级时间戳的方法
一、安装Whisper环境(Ubuntu 20.04 x64 LTS + Nvidia RTX 3090测试验证)
基础安装流程
执行以下命令完成环境搭建:
conda create -y --name whisperpy39 python==3.9 conda activate whisperpy39 pip install git+https://github.com/openai/whisper.git sudo apt update && sudo apt install ffmpeg # 基础转录测试 whisper recording.wav # 使用大模型提升转录精度 whisper recording.wav --model large
Nvidia RTX 3090专属配置
激活whisperpy39环境后,额外执行以下命令配置CUDA加速:
pip install -f https://download.pytorch.org/whl/torch_stable.html conda install pytorch==1.10.1 torchvision torchaudio cudatoolkit=11.0 -c pytorch
二、获取单词级时间戳
通过Whisper的Python API,只需指定word_timestamps=True参数,即可获取包含单词级时间戳的转录结果。示例代码如下:
import whisper # 加载指定模型,可根据需求选择tiny/base/small/medium/large model = whisper.load_model("large") # 开启单词时间戳功能并执行转录 result = model.transcribe("recording.wav", word_timestamps=True) # 解析并输出单词及对应时间戳 for segment in result["segments"]: print(f"段落时间: [{segment['start']:.2f}s - {segment['end']:.2f}s]") print(f"段落内容: {segment['text']}") print("单词明细:") for word_info in segment["words"]: print(f" '{word_info['word']}' | 起始: {word_info['start']:.2f}s | 结束: {word_info['end']:.2f}s")
执行代码后,每个单词的起始、结束时间戳会被精准输出,方便后续处理。
内容的提问来源于stack exchange,提问作者Franck Dernoncourt
相关产品推荐
相关产品推荐

