You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Distil-Whisper时return_timestamps=True导致时间戳每30秒重置的问题求助

Distil-Whisper时return_timestamps=True导致时间戳每30秒重置的问题求助

问题背景

我正在用Transformers pipeline调用Distil-Whisper做语音识别,当设置return_timestamps=True时,时间戳会每30秒重置为0,而不是在整个音频文件中持续递增。

我的代码

pipe = pipeline(
    "automatic-speech-recognition",
    model=model,
    tokenizer=processor.tokenizer,
    feature_extractor=processor.feature_extractor,
    max_new_tokens=128,
    torch_dtype=torch_dtype,
    device=device,
    return_timestamps=True,
)

result = pipe("audio.mp4")

当前输出

输出里的时间戳是这样的:

{'chunks': [
    {'text': 'First segment', 'timestamp': (0.0, 5.2)},
    {'text': 'Second segment', 'timestamp': (5.2, 12.8)},
    {'text': 'Later segment', 'timestamp': (28.4, 30.0)},
    {'text': 'Should be ~35s but shows', 'timestamp': (0.0, 4.6)},  # 这里重置了!
    ...
]}

预期行为

我期望时间戳能跨过30秒继续递增,比如:

{'chunks': [
    {'text': 'First segment', 'timestamp': (0.0, 5.2)},
    {'text': 'Second segment', 'timestamp': (5.2, 12.8)},
    {'text': 'Later segment', 'timestamp': (28.4, 30.0)},
    {'text': 'Continues properly', 'timestamp': (30.0, 34.6)},  # 应该正常延续
    ...
]}

环境信息

  • Python 3.10
  • transformers 4.36.2
  • torch 2.1.2
  • 模型:distil-whisper-large-v3

请问该如何修复这个时间戳重置的问题?有没有办法让时间戳在整个音频文件中持续递增?

备注:内容来源于stack exchange,提问作者Martin Zhu

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.14 14:43:11