如何实现带条件的递归函数?附文本匹配场景问题排查
实现基于条件的递归文本匹配分段方案
我明白你的需求了——要把长文本拆成能匹配字典值的片段,每匹配到一段就输出对应的元数据+文本,没匹配到就砍掉最后一个词重试,直到处理完所有内容。你的现有代码逻辑绕了点,没处理好剩余文本的递归匹配,我给你用递归思路重新实现,完美贴合你的需求:
完整代码实现
def find_matching_segment(line, table): # 从最长的文本片段开始尝试匹配,优先匹配最长的有效片段 words = line.strip().split() # 从完整文本逐步缩短到单个词 for i in range(len(words), 0, -1): candidate = " ".join(words[:i]) # 遍历字典检查是否是某个值的子串 for metadata, text_segment in table.items(): if candidate in text_segment: # 返回匹配到的元数据、片段,以及剩余未处理的文本 return metadata, candidate, " ".join(words[i:]) # 如果没有任何匹配,返回空值 return None, None, None def process_line_recursive(line, table, output_file): # 规整文本:去除首尾空白,合并多余空格 cleaned_line = " ".join(line.strip().split()) if not cleaned_line: return # 空文本直接结束递归 # 查找当前文本中可匹配的最长片段 metadata, matched_segment, remaining_line = find_matching_segment(cleaned_line, table) if metadata: # 写入匹配结果 output_file.write(f"{metadata}{matched_segment}\n") # 递归处理剩余的文本内容 process_line_recursive(remaining_line, table, output_file) else: # 如果完全匹配不到,可根据需求调整(比如跳过或写入未匹配内容) pass # ---------------------- 主程序 ---------------------- # 你的示例字典数据 table = { '[Base Font : IOHLGA+Trebuchet, Font Size : 3.768, Font Weight : 0.0]': 'Additions based on tax positions related to the They believe that it is reasonably possible that approximately $40 million of its unrecognized tax benefits may be recognized by the end of 2019 as a result of a lapse of the statute of limitations or resolution with the tax authorities.', '[Base Font : IOFOEO+Imago-Book, Font Size : 6.84, Font Weight : 0.0]': 'Additions based on tax positions related to prior' } # 处理文件读写 with open("myfile1.txt","r", encoding='utf-8') as f, open("myfile2.txt","w") as f1: for line in f: if line.strip(): process_line_recursive(line, table, f1) else: f1.write(line)
代码逻辑说明
find_matching_segment函数:负责从当前文本中找到最长的、能匹配字典值子串的片段。从完整文本开始逐步缩短,确保优先匹配最长的有效内容,避免拆分过细。process_line_recursive递归函数:核心处理逻辑,先规整文本,找到匹配片段后写入结果,再递归处理剩余的文本,直到所有内容都处理完毕。- 文件处理部分:遍历输入文件的每一行,非空行交给递归函数处理,空行直接写入输出文件,保留原格式。
为什么这能解决你的问题?
你的原有代码没有正确处理「匹配到片段后继续处理剩余文本」的逻辑,导致只能输出一段内容;递归方案天然适合这种“处理一部分,再处理剩余部分”的场景。同时从最长片段开始匹配,确保不会把本来可以匹配的长片段拆成短的,完全符合你预期的输出格式,还规整了文本的空格处理,避免因多余空格导致匹配失败的问题。
内容的提问来源于stack exchange,提问作者Crusader
相关产品推荐
相关产品推荐

