You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何实现带条件的递归函数?附文本匹配场景问题排查

实现基于条件的递归文本匹配分段方案

我明白你的需求了——要把长文本拆成能匹配字典值的片段,每匹配到一段就输出对应的元数据+文本,没匹配到就砍掉最后一个词重试,直到处理完所有内容。你的现有代码逻辑绕了点,没处理好剩余文本的递归匹配,我给你用递归思路重新实现,完美贴合你的需求:

完整代码实现

def find_matching_segment(line, table):
    # 从最长的文本片段开始尝试匹配,优先匹配最长的有效片段
    words = line.strip().split()
    # 从完整文本逐步缩短到单个词
    for i in range(len(words), 0, -1):
        candidate = " ".join(words[:i])
        # 遍历字典检查是否是某个值的子串
        for metadata, text_segment in table.items():
            if candidate in text_segment:
                # 返回匹配到的元数据、片段,以及剩余未处理的文本
                return metadata, candidate, " ".join(words[i:])
    # 如果没有任何匹配,返回空值
    return None, None, None

def process_line_recursive(line, table, output_file):
    # 规整文本:去除首尾空白,合并多余空格
    cleaned_line = " ".join(line.strip().split())
    if not cleaned_line:
        return  # 空文本直接结束递归
    
    # 查找当前文本中可匹配的最长片段
    metadata, matched_segment, remaining_line = find_matching_segment(cleaned_line, table)
    if metadata:
        # 写入匹配结果
        output_file.write(f"{metadata}{matched_segment}\n")
        # 递归处理剩余的文本内容
        process_line_recursive(remaining_line, table, output_file)
    else:
        # 如果完全匹配不到,可根据需求调整(比如跳过或写入未匹配内容)
        pass

# ---------------------- 主程序 ----------------------
# 你的示例字典数据
table = {
    '[Base Font : IOHLGA+Trebuchet, Font Size : 3.768, Font Weight : 0.0]': 'Additions based on tax positions related to the They believe that it is reasonably possible that approximately $40 million of its unrecognized tax benefits may be recognized by the end of 2019 as a result of a lapse of the statute of limitations or resolution with the tax authorities.',
    '[Base Font : IOFOEO+Imago-Book, Font Size : 6.84, Font Weight : 0.0]': 'Additions based on tax positions related to prior'
}

# 处理文件读写
with open("myfile1.txt","r", encoding='utf-8') as f, open("myfile2.txt","w") as f1:
    for line in f:
        if line.strip():
            process_line_recursive(line, table, f1)
        else:
            f1.write(line)

代码逻辑说明

  • find_matching_segment函数:负责从当前文本中找到最长的、能匹配字典值子串的片段。从完整文本开始逐步缩短,确保优先匹配最长的有效内容,避免拆分过细。
  • process_line_recursive递归函数:核心处理逻辑,先规整文本,找到匹配片段后写入结果,再递归处理剩余的文本,直到所有内容都处理完毕。
  • 文件处理部分:遍历输入文件的每一行,非空行交给递归函数处理,空行直接写入输出文件,保留原格式。

为什么这能解决你的问题?

你的原有代码没有正确处理「匹配到片段后继续处理剩余文本」的逻辑,导致只能输出一段内容;递归方案天然适合这种“处理一部分,再处理剩余部分”的场景。同时从最长片段开始匹配,确保不会把本来可以匹配的长片段拆成短的,完全符合你预期的输出格式,还规整了文本的空格处理,避免因多余空格导致匹配失败的问题。

内容的提问来源于stack exchange,提问作者Crusader

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.09 11:28:13