You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中Speech Recognition模块能否获取特定单词的音频时长/样本?

实现单词级音频时间戳提取与音频片段截取

SpeechRecognition模块本身不支持返回单词级别的时间戳,要实现你的需求,推荐使用OpenAI的Whisper工具——它能直接输出每个识别单词的起止时间,再结合spaCy完成词汇匹配,最后提取对应音频片段。具体步骤如下:

1. 安装依赖工具

先安装所需的Python库:

pip install openai-whisper spacy pydub
python -m spacy download en_core_web_sm

2. 用Whisper识别音频并获取词级时间戳

Whisper的转录功能支持开启词级时间戳,能精准记录每个单词的发音时段:

import whisper

# 加载Whisper模型(可选base/small/medium等,模型越大精度越高、速度越慢)
model = whisper.load_model("base")
# 转录音频,启用词级时间戳
result = model.transcribe("your_audio_file.mp3", word_timestamps=True)

# 整理所有带时间戳的单词数据
word_time_data = []
for segment in result["segments"]:
    for word_info in segment["words"]:
        word_time_data.append({
            "word": word_info["word"].strip(),
            "start": word_info["start"],
            "end": word_info["end"]
        })

3. 结合spaCy匹配目标单词的时间戳

用spaCy处理转录文本后,匹配你要找的目标单词(比如"orange"):

import spacy

# 加载spaCy英文模型
nlp = spacy.load("en_core_web_sm")
full_transcript = result["text"]
doc = nlp(full_transcript)

# 查找目标单词的所有时间戳
target_word = "orange"
matched_timestamps = [
    item for item in word_time_data
    if item["word"].lower() == target_word.lower()
]

if matched_timestamps:
    for ts in matched_timestamps:
        print(f"单词'{target_word}'的发音时段:{ts['start']:.2f}秒 至 {ts['end']:.2f}秒")
else:
    print(f"未在音频中识别到单词'{target_word}'")

4. 提取对应单词的音频样本

用pydub裁剪音频,导出目标单词对应的片段:

from pydub import AudioSegment

# 加载原音频文件
original_audio = AudioSegment.from_file("your_audio_file.mp3")

# 遍历匹配到的时间戳,裁剪并保存音频片段
for idx, ts in enumerate(matched_timestamps):
    # 转换秒为毫秒(pydub以毫秒为单位)
    start_ms = ts["start"] * 1000
    end_ms = ts["end"] * 1000
    # 裁剪音频
    word_segment = original_audio[start_ms:end_ms]
    # 导出为独立文件
    word_segment.export(f"{target_word}_segment_{idx+1}.mp3", format="mp3")
    print(f"已保存'{target_word}'的音频片段:{target_word}_segment_{idx+1}.mp3")

注意事项

  • Whisper的word_timestamps=True是开启词级时间戳的核心参数,大模型(如medium)的时间戳精度优于小模型;
  • 匹配单词时统一转小写,避免因大小写差异遗漏结果;
  • 如果处理非英语音频,需加载对应语言的Whisper模型和spaCy模型。

内容的提问来源于stack exchange,提问作者user16335562

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.16 07:30:54