You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

运行lazykh唇读项目scheduler.py报ValueError: substring not found错误

lazykh项目scheduler.py substring not found报错修复

报错基础信息

  • 触发场景:运行carykh/lazykh开源唇读项目第三步流程时触发运行时错误
  • 异常位置:本地路径C:\Users\User\Desktop\lazykh-main\code\scheduler.py第93行
  • 触发执行语句:OS_nextIndex = originalScript.index(wordString,OS_IndexAt)+len(wordString)
  • 抛出异常:ValueError: substring not found

根本诱因

原代码逻辑默认语音识别返回的单词序列和原始txt脚本文本严格逐字顺序、逐字符完全匹配,未做任何容错处理,以下任意场景都会触发子串查找失败:

  • 文本大小写不匹配:原始脚本单词大小写和语音转写返回的wordString不一致,比如脚本写Hello、转写返回hello
  • 标点粘连:原始脚本中单词和后缀标点直接相连,比如脚本写world!、转写返回纯单词world
  • 索引偏移:脚本内的<...>控制标签、多余换行/空格、标签嵌套/交叉场景下,原有标签跳过逻辑计算OS_IndexAt错误,查找起始位置已经越过目标单词所在区间
  • 转写偏差:语音识别返回not-found-in-audio标记词,或是转写结果和原始脚本存在缩写/全写差异(比如脚本写cannot、转写返回can't)

可落地修复方案

1. 替换硬查找逻辑,增加多层容错

将原代码第90~96行的硬匹配逻辑:

wordString = word["word"]
timeStart = word["start"]
OS_nextIndex = originalScript.index(wordString,OS_IndexAt)+len(wordString)
if "<" in originalScript[OS_IndexAt:]:
    tagStart = originalScript.index("<",OS_IndexAt)
    tagEnd = originalScript.index(">",OS_IndexAt)
    if OS_nextIndex > tagStart and tagEnd >= OS_nextIndex:
        OS_nextIndex = originalScript.index(wordString,tagEnd)+len(wordString)

替换为以下容错版本,覆盖大小写匹配、标签跳过、标点剥离、兜底防崩溃逻辑:

wordString = word["word"].lower()
timeStart = word["start"]
OS_nextIndex = -1
search_pos = OS_IndexAt
# 先跳过所有<...>标签块,在标签间隙文本中查找目标单词
while "<" in originalScript[search_pos:]:
    tagStart = originalScript.index("<", search_pos)
    tagEnd = originalScript.index(">", search_pos)
    pre_tag_text = originalScript[search_pos:tagStart].lower()
    if wordString in pre_tag_text:
        word_rel_pos = pre_tag_text.index(wordString)
        OS_nextIndex = search_pos + word_rel_pos + len(wordString)
        break
    search_pos = tagEnd + 1
# 标签区域未命中,在剩余文本中做去标点模糊匹配
if OS_nextIndex == -1:
    import re
    remain_text = originalScript[search_pos:].lower()
    clean_punct = re.compile(r'[,\.;:!\?\s"\'()\-]+')
    target_word_clean = clean_punct.sub('', wordString)
    text_cursor = 0
    while text_cursor < len(remain_text):
        # 跳过非字母字符
        while text_cursor < len(remain_text) and not remain_text[text_cursor].isalpha():
            text_cursor += 1
        if text_cursor >= len(remain_text):
            break
        # 提取当前位置连续单词
        word_end = text_cursor
        while word_end < len(remain_text) and (remain_text[word_end].isalpha() or remain_text[word_end] == "'"):
            word_end += 1
        cur_word_clean = clean_punct.sub('', remain_text[text_cursor:word_end])
        if cur_word_clean == target_word_clean:
            OS_nextIndex = search_pos + word_end
            break
        text_cursor = word_end
# 极端兜底:匹配失败时自动推进索引,避免程序直接崩溃
if OS_nextIndex == -1:
    print(f"警告:未匹配到单词{word['word']},已自动跳过")
    OS_nextIndex = OS_IndexAt + len(wordString)

2. 前置优化减少匹配冲突

  • 统一文本大小写:读取原始脚本时直接统一转小写,将原代码中originalScript = f.read()修改为originalScript = f.read().lower()
  • 跳过无效词项:在循环开头增加异常词判断,直接跳过语音识别未命中的词,在if "start" not in word: continue语句下新增一行:
    if word.get("case") == "not-found-in-audio":
        continue
    
  • 输入脚本预处理:运行前检查输入txt文件,确保所有<...>控制标签正常闭合,不要在普通文本中使用半角</>符号,删除多余的全角特殊字符。

内容的提问来源于stack exchange,提问作者bease

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.28 07:39:20