文件正则匹配追加内容及标识替换技术实现咨询
问题描述
需要实现两个核心功能:
- 遍历目标
.txt文件,匹配到格式为===r(xxxx).(xxxx).(xxxx)的行时,给后续所有包含remark的行末尾追加该r(xxxx).(xxxx).(xxxx)标识,直到遇到包含remark =-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=的行,重复此逻辑至文件结束。 - 将所有
r(xxxx)标识替换为映射文件中的对应新词汇(映射文件格式为r1130,new_word)。
现有Python代码仅能匹配首个正则结果,后续处理存在问题,需优化。
示例文件片段
access-list r1999-outside-in remark =-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-= access-list r1999-outside-in remark ===r9920.4001.886 access-list r1999-outside-in remark Access from Test Network access-list r1999-outside-in extended permit tcp xxx.xxx.xxx.xxx 255.255.255.128 x.x.32.160 255.255.255.248 eq 4343 ...
预期处理效果片段
access-list r1999-outside-in remark =-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-= access-list r1999-outside-in remark ===r9920.4001.886 access-list r1999-outside-in remark Access from Test Network r9920.4001.886 access-list r1999-outside-in extended permit tcp xxx.xxx.xxx.xxx 255.255.255.128 x.x.32.160 255.255.255.248 eq 4343 ...
当前代码片段
with open("file.txt", "r") as file: # Create an empty list to store the lines lines = [] # Iterate over the lines of the file for line in file: line = line.strip() # Append the line to a list lines.append(line) # regex to pull out r(xxxx).(xxxx).(xxxx) y = re.search("(r\d{4}.\d+.\w+)", line) if y != None: print(y)
优化解决方案
核心思路
- 状态跟踪:用变量维护当前有效的标识,解决单次匹配的问题
- 分阶段处理:先完成标识追加,再统一处理映射替换
- 精准正则匹配:针对不同行类型编写独立规则,避免误匹配
完整优化代码
import re def process_remark_file(input_file, mapping_file, output_file): # 读取映射文件,构建替换字典 mapping_dict = {} with open(mapping_file, 'r', encoding='utf-8') as f: for line in f: line = line.strip() if line: old_tag, new_word = line.split(',', 1) mapping_dict[old_tag.strip()] = new_word.strip() current_tag = None processed_lines = [] # 定义正则匹配规则 split_line_re = re.compile(r'remark =-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=') tag_line_re = re.compile(r'===r(\d{4}\.\d+\.\w+)') remark_line_re = re.compile(r'remark ') with open(input_file, 'r', encoding='utf-8') as f: for line in f: original_line = line.rstrip('\n') # 保留原始行格式(除换行符) stripped_line = original_line.strip() # 遇到分割线,重置当前标识 if split_line_re.search(original_line): current_tag = None processed_lines.append(original_line) continue # 匹配到标识行,提取并保存标识 tag_match = tag_line_re.search(original_line) if tag_match: current_tag = tag_match.group(1) processed_lines.append(original_line) continue # 普通remark行,追加当前标识 if remark_line_re.search(original_line) and current_tag: # 跳过标识行和分割线的重复追加 if not tag_line_re.search(original_line) and not split_line_re.search(original_line): processed_lines.append(f"{original_line} {current_tag}") continue # 其他行直接保留 processed_lines.append(original_line) # 替换所有标识为映射词汇 final_lines = [] for line in processed_lines: replaced_line = line for old_tag, new_word in mapping_dict.items(): replaced_line = re.sub(re.escape(old_tag), new_word, replaced_line) final_lines.append(replaced_line) # 写入结果文件 with open(output_file, 'w', encoding='utf-8') as f: f.write('\n'.join(final_lines) + '\n') # 调用示例 process_remark_file("file.txt", "mapping.txt", "processed_file.txt")
关键优化点
- 状态维护:通过
current_tag变量持续跟踪当前需追加的标识,实现多段内容的循环处理 - 格式兼容性:保留原始行的缩进和空白,避免破坏文件原有格式
- 正则精准性:针对分割线、标识行、普通remark行分别编写正则,减少误匹配概率
- 映射效率:提前构建映射字典,批量替换时直接查询,提升处理速度
内容的提问来源于stack exchange,提问作者Adema_It_Man
相关产品推荐
相关产品推荐

