Azure认知服务语音SDK生成SRT字幕缺失序号问题求助
修复Azure Speech SDK Python字幕生成SRT缺失序号问题
问题描述
使用Azure认知服务语音SDK的Python字幕生成代码生成SRT格式字幕时,输出缺失字幕序号,仅显示时间轴和文本内容:
00:00:00,180 --> 00:00:03,230 Welcome to applied Mathematics course 201.
期望输出应为:
1 00:00:00,180 --> 00:00:03,230 Welcome to applied Mathematics course 201.
自行修改第74行的sequenceNumber为sequence_number后仍未解决问题。
问题根源
代码中存在两处关键问题:
- 变量名不匹配:
caption_from_speech_recognition_result函数内引用了未定义的sequenceNumber变量,正确的参数名应为sequence_number(函数定义时使用下划线命名)。 - 运行参数逻辑:若未添加
--srt参数,代码不会触发SRT格式的序号生成逻辑;若启用了--recognizing参数,序号生成条件会被跳过(因为条件判断包含not user_config["show_recognizing_results"])。
修复步骤
1. 修正变量名错误
找到caption_from_speech_recognition_result函数(约第70-78行),将其中的sequenceNumber替换为sequence_number:
def caption_from_speech_recognition_result(sequence_number : int, result : speechsdk.SpeechRecognitionResult, user_config : helper.Read_Only_Dict) -> str : caption = "" if not user_config["show_recognizing_results"] and user_config["use_sub_rip_text_caption_format"] : # 修正此处变量名 caption += str(sequence_number) + linesep caption += timestamp_from_speech_recognition_result(result = result, user_config = user_config) + linesep caption += language_from_speech_recognition_result(result = result, user_config = user_config) caption += result.text + linesep + linesep return caption
2. 确保运行参数正确
运行代码时必须添加--srt参数以启用SRT格式输出,示例命令:
python captioning.py --key 你的订阅密钥 --region 你的区域 --input 输入音频文件.wav --output 输出字幕.srt --srt
- 若不需要实时显示识别过程,请勿添加
--recognizing参数(该参数会跳过序号生成逻辑)。
3. 验证序号递增逻辑
确认recognized_handler中的序号递增逻辑正常:
def recognized_handler(e : speechsdk.SpeechRecognitionEventArgs) : if speechsdk.ResultReason.RecognizedSpeech == e.result.reason and len(e.result.text) > 0 : nonlocal sequence_number sequence_number += 1 helper.write_to_console_or_file(text = caption_from_speech_recognition_result(sequence_number = sequence_number, result = e.result, user_config = user_config), user_config = user_config)
这段代码会在每次完成语音识别后,将sequence_number加1并传入字幕生成函数,确保序号连续递增。
内容的提问来源于stack exchange,提问作者fedep11
相关产品推荐
相关产品推荐

