You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在OpenAI Whisper语音识别中获取单词级时间戳?

用OpenAI Whisper Python库获取单词级时间戳的方法

一、安装Whisper环境(Ubuntu 20.04 x64 LTS + Nvidia RTX 3090测试验证)

基础安装流程

执行以下命令完成环境搭建:

conda create -y --name whisperpy39 python==3.9
conda activate whisperpy39
pip install git+https://github.com/openai/whisper.git 
sudo apt update && sudo apt install ffmpeg
# 基础转录测试
whisper recording.wav
# 使用大模型提升转录精度
whisper recording.wav --model large

Nvidia RTX 3090专属配置

激活whisperpy39环境后,额外执行以下命令配置CUDA加速:

pip install -f https://download.pytorch.org/whl/torch_stable.html
conda install pytorch==1.10.1 torchvision torchaudio cudatoolkit=11.0 -c pytorch

二、获取单词级时间戳

通过Whisper的Python API,只需指定word_timestamps=True参数,即可获取包含单词级时间戳的转录结果。示例代码如下:

import whisper

# 加载指定模型,可根据需求选择tiny/base/small/medium/large
model = whisper.load_model("large")

# 开启单词时间戳功能并执行转录
result = model.transcribe("recording.wav", word_timestamps=True)

# 解析并输出单词及对应时间戳
for segment in result["segments"]:
    print(f"段落时间: [{segment['start']:.2f}s - {segment['end']:.2f}s]")
    print(f"段落内容: {segment['text']}")
    print("单词明细:")
    for word_info in segment["words"]:
        print(f"  '{word_info['word']}' | 起始: {word_info['start']:.2f}s | 结束: {word_info['end']:.2f}s")

执行代码后,每个单词的起始、结束时间戳会被精准输出,方便后续处理。

内容的提问来源于stack exchange,提问作者Franck Dernoncourt

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.18 15:35:26