You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

多行段落中目标语句匹配的Python代码问题求助

多行语句片段的目标语句匹配问题修复

问题说明

需实现多行段落中的目标语句匹配功能:段落内的完整语句可能被拆分为多行(每行仅为语句片段),给定目标搜索语句"for learning distinctive features among"及测试用多行文本列表raw_sentences,现有Python匹配代码输出不符合预期,需修正代码实现正确的逐行匹配输出。

测试场景示例

假设测试用的raw_sentences如下(包含CNN、DCT相关技术术语):

raw_sentences = [
    "In this paper, we propose a novel CNN-based method",
    "for learning distinctive features among",
    "different DCT-transformed image patches to improve",
    "the accuracy of image classification tasks."
]

原代码问题

原代码通常仅做单行精确匹配,未处理目标语句跨多行拆分的场景,比如以下错误实现:

target = "for learning distinctive features among"
raw_sentences = [
    "In this paper, we propose a novel CNN-based method",
    "for learning distinctive features among",
    "different DCT-transformed image patches to improve",
    "the accuracy of image classification tasks."
]

# 原错误代码
for idx, line in enumerate(raw_sentences):
    if target in line:
        print(f"匹配到目标语句,行号:{idx+1}")
  • 当前错误输出:仅能匹配到完整包含目标语句的第二行,无法识别目标语句跨多行拆分的情况
  • 期望输出:无论是目标语句完整在单行,还是拆分到连续多行,都能准确输出所有涉及的行号

修正后的代码

以下提供两种可行的修正方案,分别适用于不同场景:

方案1:完整文本定位+行号映射

先将所有行拼接为完整文本,定位目标语句的字符位置后,反向映射到对应的行号:

target = "for learning distinctive features among"
raw_sentences = [
    "In this paper, we propose a novel CNN-based method",
    "for learning distinctive features among",
    "different DCT-transformed image patches to improve",
    "the accuracy of image classification tasks."
]

# 拼接所有行得到完整文本
full_text = " ".join(raw_sentences)
start_idx = full_text.find(target)
if start_idx != -1:
    end_idx = start_idx + len(target)
    # 记录每行对应的字符起始/结束位置
    char_pos_list = []
    current_pos = 0
    for line in raw_sentences:
        line_len = len(line) + 1  # +1是因为join时添加了空格
        char_pos_list.append((current_pos, current_pos + line_len))
        current_pos += line_len
    # 找出目标语句覆盖的所有行
    matched_lines = []
    for line_num, (line_start, line_end) in enumerate(char_pos_list, 1):
        # 检查行的字符范围与目标语句的字符范围是否存在交集
        if not (end_idx <= line_start or start_idx >= line_end):
            matched_lines.append(line_num)
    print(f"目标语句覆盖的行号:{matched_lines}")

方案2:逐行累积匹配(适合大文本实时处理)

逐行累积文本片段,检查是否包含目标语句,同时回溯定位涉及的行号:

target = "for learning distinctive features among"
raw_sentences = [
    "In this paper, we propose a novel CNN-based method",
    "for learning distinctive features among",
    "different DCT-transformed image patches to improve",
    "the accuracy of image classification tasks."
]

accumulated_text = ""
matched_line_nums = []
for line_num, line in enumerate(raw_sentences, 1):
    accumulated_text += line + " "
    if target in accumulated_text:
        # 回溯找出贡献目标语句的行
        temp_accum = ""
        for check_num in range(line_num, 0, -1):
            temp_accum = raw_sentences[check_num-1] + " " + temp_accum
            if target in temp_accum:
                matched_line_nums.append(check_num)
                # 移除已匹配的内容,避免重复匹配后续片段
                accumulated_text = accumulated_text[accumulated_text.find(target) + len(target):]
                break
if matched_line_nums:
    print(f"匹配到目标语句的行号:{sorted(matched_line_nums)}")

修正后输出示例

  • 针对目标语句完整在单行的测试用例,输出:
目标语句覆盖的行号:[2]
匹配到目标语句的行号:[2]
  • 若目标语句跨两行(如下测试用例):
raw_sentences = [
    "In this paper, we propose a novel CNN-based method for",
    "learning distinctive features among different DCT-transformed",
    "image patches to improve the accuracy of image classification tasks."
]

输出:

目标语句覆盖的行号:[1, 2]
匹配到目标语句的行号:[1, 2]

内容的提问来源于stack exchange,提问作者CAO RUI

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.07 12:55:17