多行段落中目标语句匹配的Python代码问题求助
多行语句片段的目标语句匹配问题修复
问题说明
需实现多行段落中的目标语句匹配功能:段落内的完整语句可能被拆分为多行(每行仅为语句片段),给定目标搜索语句"for learning distinctive features among"及测试用多行文本列表raw_sentences,现有Python匹配代码输出不符合预期,需修正代码实现正确的逐行匹配输出。
测试场景示例
假设测试用的raw_sentences如下(包含CNN、DCT相关技术术语):
raw_sentences = [ "In this paper, we propose a novel CNN-based method", "for learning distinctive features among", "different DCT-transformed image patches to improve", "the accuracy of image classification tasks." ]
原代码问题
原代码通常仅做单行精确匹配,未处理目标语句跨多行拆分的场景,比如以下错误实现:
target = "for learning distinctive features among" raw_sentences = [ "In this paper, we propose a novel CNN-based method", "for learning distinctive features among", "different DCT-transformed image patches to improve", "the accuracy of image classification tasks." ] # 原错误代码 for idx, line in enumerate(raw_sentences): if target in line: print(f"匹配到目标语句,行号:{idx+1}")
- 当前错误输出:仅能匹配到完整包含目标语句的第二行,无法识别目标语句跨多行拆分的情况
- 期望输出:无论是目标语句完整在单行,还是拆分到连续多行,都能准确输出所有涉及的行号
修正后的代码
以下提供两种可行的修正方案,分别适用于不同场景:
方案1:完整文本定位+行号映射
先将所有行拼接为完整文本,定位目标语句的字符位置后,反向映射到对应的行号:
target = "for learning distinctive features among" raw_sentences = [ "In this paper, we propose a novel CNN-based method", "for learning distinctive features among", "different DCT-transformed image patches to improve", "the accuracy of image classification tasks." ] # 拼接所有行得到完整文本 full_text = " ".join(raw_sentences) start_idx = full_text.find(target) if start_idx != -1: end_idx = start_idx + len(target) # 记录每行对应的字符起始/结束位置 char_pos_list = [] current_pos = 0 for line in raw_sentences: line_len = len(line) + 1 # +1是因为join时添加了空格 char_pos_list.append((current_pos, current_pos + line_len)) current_pos += line_len # 找出目标语句覆盖的所有行 matched_lines = [] for line_num, (line_start, line_end) in enumerate(char_pos_list, 1): # 检查行的字符范围与目标语句的字符范围是否存在交集 if not (end_idx <= line_start or start_idx >= line_end): matched_lines.append(line_num) print(f"目标语句覆盖的行号:{matched_lines}")
方案2:逐行累积匹配(适合大文本实时处理)
逐行累积文本片段,检查是否包含目标语句,同时回溯定位涉及的行号:
target = "for learning distinctive features among" raw_sentences = [ "In this paper, we propose a novel CNN-based method", "for learning distinctive features among", "different DCT-transformed image patches to improve", "the accuracy of image classification tasks." ] accumulated_text = "" matched_line_nums = [] for line_num, line in enumerate(raw_sentences, 1): accumulated_text += line + " " if target in accumulated_text: # 回溯找出贡献目标语句的行 temp_accum = "" for check_num in range(line_num, 0, -1): temp_accum = raw_sentences[check_num-1] + " " + temp_accum if target in temp_accum: matched_line_nums.append(check_num) # 移除已匹配的内容,避免重复匹配后续片段 accumulated_text = accumulated_text[accumulated_text.find(target) + len(target):] break if matched_line_nums: print(f"匹配到目标语句的行号:{sorted(matched_line_nums)}")
修正后输出示例
- 针对目标语句完整在单行的测试用例,输出:
目标语句覆盖的行号:[2] 匹配到目标语句的行号:[2]
- 若目标语句跨两行(如下测试用例):
raw_sentences = [ "In this paper, we propose a novel CNN-based method for", "learning distinctive features among different DCT-transformed", "image patches to improve the accuracy of image classification tasks." ]
输出:
目标语句覆盖的行号:[1, 2] 匹配到目标语句的行号:[1, 2]
内容的提问来源于stack exchange,提问作者CAO RUI
相关产品推荐
相关产品推荐

