You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Azure认知服务语音SDK生成SRT字幕缺失序号问题求助

修复Azure Speech SDK Python字幕生成SRT缺失序号问题

问题描述

使用Azure认知服务语音SDK的Python字幕生成代码生成SRT格式字幕时,输出缺失字幕序号,仅显示时间轴和文本内容:

00:00:00,180 --> 00:00:03,230
Welcome to applied Mathematics course 201.

期望输出应为:

1
00:00:00,180 --> 00:00:03,230
Welcome to applied Mathematics course 201.

自行修改第74行的sequenceNumber为sequence_number后仍未解决问题。

问题根源

代码中存在两处关键问题:

  1. 变量名不匹配:caption_from_speech_recognition_result函数内引用了未定义的sequenceNumber变量,正确的参数名应为sequence_number(函数定义时使用下划线命名)。
  2. 运行参数逻辑:若未添加--srt参数,代码不会触发SRT格式的序号生成逻辑;若启用了--recognizing参数,序号生成条件会被跳过(因为条件判断包含not user_config["show_recognizing_results"])。

修复步骤

1. 修正变量名错误

找到caption_from_speech_recognition_result函数(约第70-78行),将其中的sequenceNumber替换为sequence_number:

def caption_from_speech_recognition_result(sequence_number : int, result : speechsdk.SpeechRecognitionResult, user_config : helper.Read_Only_Dict) -> str :
    caption = ""
    if not user_config["show_recognizing_results"] and user_config["use_sub_rip_text_caption_format"] :
        # 修正此处变量名
        caption += str(sequence_number) + linesep
    caption += timestamp_from_speech_recognition_result(result = result, user_config = user_config) + linesep
    caption += language_from_speech_recognition_result(result = result, user_config = user_config)
    caption += result.text + linesep + linesep
    return caption

2. 确保运行参数正确

运行代码时必须添加--srt参数以启用SRT格式输出,示例命令:

python captioning.py --key 你的订阅密钥 --region 你的区域 --input 输入音频文件.wav --output 输出字幕.srt --srt
  • 若不需要实时显示识别过程,请勿添加--recognizing参数(该参数会跳过序号生成逻辑)。

3. 验证序号递增逻辑

确认recognized_handler中的序号递增逻辑正常:

def recognized_handler(e : speechsdk.SpeechRecognitionEventArgs) :
    if speechsdk.ResultReason.RecognizedSpeech == e.result.reason and len(e.result.text) > 0 :
        nonlocal sequence_number
        sequence_number += 1
        helper.write_to_console_or_file(text = caption_from_speech_recognition_result(sequence_number = sequence_number, result = e.result, user_config = user_config), user_config = user_config)

这段代码会在每次完成语音识别后,将sequence_number加1并传入字幕生成函数,确保序号连续递增。

内容的提问来源于stack exchange,提问作者fedep11

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.19 03:10:25