使用Distil-Whisper时return_timestamps=True导致时间戳每30秒重置的问题求助
Distil-Whisper时return_timestamps=True导致时间戳每30秒重置的问题求助
问题背景
我正在用Transformers pipeline调用Distil-Whisper做语音识别,当设置return_timestamps=True时,时间戳会每30秒重置为0,而不是在整个音频文件中持续递增。
我的代码
pipe = pipeline( "automatic-speech-recognition", model=model, tokenizer=processor.tokenizer, feature_extractor=processor.feature_extractor, max_new_tokens=128, torch_dtype=torch_dtype, device=device, return_timestamps=True, ) result = pipe("audio.mp4")
当前输出
输出里的时间戳是这样的:
{'chunks': [ {'text': 'First segment', 'timestamp': (0.0, 5.2)}, {'text': 'Second segment', 'timestamp': (5.2, 12.8)}, {'text': 'Later segment', 'timestamp': (28.4, 30.0)}, {'text': 'Should be ~35s but shows', 'timestamp': (0.0, 4.6)}, # 这里重置了! ... ]}
预期行为
我期望时间戳能跨过30秒继续递增,比如:
{'chunks': [ {'text': 'First segment', 'timestamp': (0.0, 5.2)}, {'text': 'Second segment', 'timestamp': (5.2, 12.8)}, {'text': 'Later segment', 'timestamp': (28.4, 30.0)}, {'text': 'Continues properly', 'timestamp': (30.0, 34.6)}, # 应该正常延续 ... ]}
环境信息
- Python 3.10
- transformers 4.36.2
- torch 2.1.2
- 模型:distil-whisper-large-v3
请问该如何修复这个时间戳重置的问题?有没有办法让时间戳在整个音频文件中持续递增?
备注:内容来源于stack exchange,提问作者Martin Zhu
相关产品推荐
相关产品推荐

