You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中如何搜索文件内多行字符串并获取起止行列信息?

问题描述

需要在Python中搜索文件内的多行字符串,匹配成功后获取以下位置元数据:

  • 起始行号
  • 结束行号
  • 起始列号(模式首个字符所在行的索引)
  • 结束列号(模式最后一个字符所在行的最后索引)

例如要匹配的多行模式为:

pattern = """b'0100000001685c7c35aabe690cc99f947a8172ad075d4401448a212b9f26607d6ec5530915010000006a4730'
           b'440220337117278ee2fc7ae222ec1547b3a40fa39a05f91c1e19db60060541c4b3d6e4022020188e1d5d843c'"""

预期匹配结果为:

  • start_line: 2
  • end_line: 3
  • start_column: 23
  • end_column: 114

尝试直接用re模块搜索返回None,代码如下:

import re

pattern = """b'0100000001685c7c35aabe690cc99f947a8172ad075d4401448a212b9f26607d6ec5530915010000006a4730'
           b'440220337117278ee2fc7ae222ec1547b3a40fa39a05f91c1e19db60060541c4b3d6e4022020188e1d5d843c'"""


with open("test.py") as f:
    content = f.read()
    print(re.search(pattern, content))

单行匹配的代码能获取位置信息,但无法处理多行场景:

with open("test.py") as f:
    data = f.read()
    for n, line in enumerate(data):
        match_index = line.find(pattern)
        if match_index != -1:
            print("Start Line:", n + 1)
            print("End Line", n + 1)
            print("Start Column:", match_index)
            print("End Column:", match_index + len(pattern) + 1)
            break

求解决办法:如何在Python中匹配文件内的多行字符串并获取匹配位置的元数据?


解决方案

1. 修复正则匹配失败问题

直接用re.search返回None的核心原因:

  • 模式中的换行、缩进和文件内目标内容的格式不一致
  • 正则默认不匹配换行符,需启用re.DOTALL参数让.可以匹配换行

如果模式中的缩进不固定,可将固定缩进替换为\s+匹配任意空白;若缩进固定,则确保模式和文件内容的空白完全一致。

2. 获取位置元数据的实现代码

通过正则匹配到内容后,利用匹配结果的起始/结束索引,结合文件的行索引计算出行号和列号:

import re

# 定义多行模式,用\s*匹配两行之间的任意空白(包括换行和缩进)
pattern = r"""b'0100000001685c7c35aabe690cc99f947a8172ad075d4401448a212b9f26607d6ec5530915010000006a4730'\s*
           b'440220337117278ee2fc7ae222ec1547b3a40fa39a05f91c1e19db60060541c4b3d6e4022020188e1d5d843c'"""

with open("test.py", "r") as f:
    content = f.read()
    # 保留换行符拆分每行,确保索引计算准确
    lines = content.splitlines(keepends=True)

# 启用DOTALL模式匹配多行内容
match = re.search(pattern, content, flags=re.DOTALL)
if match:
    start_idx = match.start()
    end_idx = match.end() - 1  # 取匹配内容最后一个字符的索引

    # 计算起始行和列
    current_pos = 0
    start_line, start_col = 0, 0
    for line_num, line in enumerate(lines):
        line_length = len(line)
        if current_pos + line_length > start_idx:
            start_line = line_num + 1  # 行号从1开始计数
            start_col = start_idx - current_pos
            break
        current_pos += line_length

    # 计算结束行和列
    current_pos = 0
    end_line, end_col = 0, 0
    for line_num, line in enumerate(lines):
        line_length = len(line)
        if current_pos + line_length > end_idx:
            end_line = line_num + 1
            end_col = end_idx - current_pos
            break
        current_pos += line_length

    # 输出结果
    print(f"start_line: {start_line}")
    print(f"end_line: {end_line}")
    print(f"start_column: {start_col}")
    print(f"end_column: {end_col}")
else:
    print("未找到匹配内容")

代码说明

  • splitlines(keepends=True):保留每行的换行符,避免因换行符丢失导致索引计算错误
  • 遍历每行累加长度,定位匹配索引对应的行号和列号
  • re.DOTALL:确保正则可以跨多行匹配内容
  • 若模式中的空白不固定,用\s+或\s*代替固定缩进,提升匹配灵活性

内容的提问来源于stack exchange,提问作者Setu Kumar Basak

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.10 11:30:50