Python中如何搜索文件内多行字符串并获取起止行列信息?
问题描述
需要在Python中搜索文件内的多行字符串,匹配成功后获取以下位置元数据:
- 起始行号
- 结束行号
- 起始列号(模式首个字符所在行的索引)
- 结束列号(模式最后一个字符所在行的最后索引)
例如要匹配的多行模式为:
pattern = """b'0100000001685c7c35aabe690cc99f947a8172ad075d4401448a212b9f26607d6ec5530915010000006a4730' b'440220337117278ee2fc7ae222ec1547b3a40fa39a05f91c1e19db60060541c4b3d6e4022020188e1d5d843c'"""
预期匹配结果为:
- start_line: 2
- end_line: 3
- start_column: 23
- end_column: 114
尝试直接用re模块搜索返回None,代码如下:
import re pattern = """b'0100000001685c7c35aabe690cc99f947a8172ad075d4401448a212b9f26607d6ec5530915010000006a4730' b'440220337117278ee2fc7ae222ec1547b3a40fa39a05f91c1e19db60060541c4b3d6e4022020188e1d5d843c'""" with open("test.py") as f: content = f.read() print(re.search(pattern, content))
单行匹配的代码能获取位置信息,但无法处理多行场景:
with open("test.py") as f: data = f.read() for n, line in enumerate(data): match_index = line.find(pattern) if match_index != -1: print("Start Line:", n + 1) print("End Line", n + 1) print("Start Column:", match_index) print("End Column:", match_index + len(pattern) + 1) break
求解决办法:如何在Python中匹配文件内的多行字符串并获取匹配位置的元数据?
解决方案
1. 修复正则匹配失败问题
直接用re.search返回None的核心原因:
- 模式中的换行、缩进和文件内目标内容的格式不一致
- 正则默认不匹配换行符,需启用
re.DOTALL参数让.可以匹配换行
如果模式中的缩进不固定,可将固定缩进替换为\s+匹配任意空白;若缩进固定,则确保模式和文件内容的空白完全一致。
2. 获取位置元数据的实现代码
通过正则匹配到内容后,利用匹配结果的起始/结束索引,结合文件的行索引计算出行号和列号:
import re # 定义多行模式,用\s*匹配两行之间的任意空白(包括换行和缩进) pattern = r"""b'0100000001685c7c35aabe690cc99f947a8172ad075d4401448a212b9f26607d6ec5530915010000006a4730'\s* b'440220337117278ee2fc7ae222ec1547b3a40fa39a05f91c1e19db60060541c4b3d6e4022020188e1d5d843c'""" with open("test.py", "r") as f: content = f.read() # 保留换行符拆分每行,确保索引计算准确 lines = content.splitlines(keepends=True) # 启用DOTALL模式匹配多行内容 match = re.search(pattern, content, flags=re.DOTALL) if match: start_idx = match.start() end_idx = match.end() - 1 # 取匹配内容最后一个字符的索引 # 计算起始行和列 current_pos = 0 start_line, start_col = 0, 0 for line_num, line in enumerate(lines): line_length = len(line) if current_pos + line_length > start_idx: start_line = line_num + 1 # 行号从1开始计数 start_col = start_idx - current_pos break current_pos += line_length # 计算结束行和列 current_pos = 0 end_line, end_col = 0, 0 for line_num, line in enumerate(lines): line_length = len(line) if current_pos + line_length > end_idx: end_line = line_num + 1 end_col = end_idx - current_pos break current_pos += line_length # 输出结果 print(f"start_line: {start_line}") print(f"end_line: {end_line}") print(f"start_column: {start_col}") print(f"end_column: {end_col}") else: print("未找到匹配内容")
代码说明
splitlines(keepends=True):保留每行的换行符,避免因换行符丢失导致索引计算错误- 遍历每行累加长度,定位匹配索引对应的行号和列号
re.DOTALL:确保正则可以跨多行匹配内容- 若模式中的空白不固定,用
\s+或\s*代替固定缩进,提升匹配灵活性
内容的提问来源于stack exchange,提问作者Setu Kumar Basak
相关产品推荐
相关产品推荐

