You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

文件正则匹配追加内容及标识替换技术实现咨询

问题描述

需要实现两个核心功能:

  1. 遍历目标.txt文件,匹配到格式为===r(xxxx).(xxxx).(xxxx)的行时,给后续所有包含remark的行末尾追加该r(xxxx).(xxxx).(xxxx)标识,直到遇到包含remark =-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=的行,重复此逻辑至文件结束。
  2. 将所有r(xxxx)标识替换为映射文件中的对应新词汇(映射文件格式为r1130,new_word)。

现有Python代码仅能匹配首个正则结果,后续处理存在问题,需优化。

示例文件片段

access-list r1999-outside-in remark =-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=
access-list r1999-outside-in remark ===r9920.4001.886
access-list r1999-outside-in remark Access from Test Network
access-list r1999-outside-in extended permit tcp xxx.xxx.xxx.xxx 255.255.255.128 x.x.32.160 255.255.255.248 eq 4343 
...

预期处理效果片段

access-list r1999-outside-in remark =-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=
access-list r1999-outside-in remark ===r9920.4001.886
access-list r1999-outside-in remark Access from Test Network r9920.4001.886
access-list r1999-outside-in extended permit tcp xxx.xxx.xxx.xxx 255.255.255.128 x.x.32.160 255.255.255.248 eq 4343 
...

当前代码片段

with open("file.txt", "r") as file:
# Create an empty list to store the lines
    lines = []

    # Iterate over the lines of the file
    for line in file:
        line = line.strip()
        # Append the line to a list
        lines.append(line)
        # regex to pull out r(xxxx).(xxxx).(xxxx)
        y = re.search("(r\d{4}.\d+.\w+)", line)
        if y != None:
            print(y)

优化解决方案

核心思路

  1. 状态跟踪:用变量维护当前有效的标识,解决单次匹配的问题
  2. 分阶段处理:先完成标识追加,再统一处理映射替换
  3. 精准正则匹配:针对不同行类型编写独立规则,避免误匹配

完整优化代码

import re

def process_remark_file(input_file, mapping_file, output_file):
    # 读取映射文件,构建替换字典
    mapping_dict = {}
    with open(mapping_file, 'r', encoding='utf-8') as f:
        for line in f:
            line = line.strip()
            if line:
                old_tag, new_word = line.split(',', 1)
                mapping_dict[old_tag.strip()] = new_word.strip()

    current_tag = None
    processed_lines = []

    # 定义正则匹配规则
    split_line_re = re.compile(r'remark =-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=')
    tag_line_re = re.compile(r'===r(\d{4}\.\d+\.\w+)')
    remark_line_re = re.compile(r'remark ')

    with open(input_file, 'r', encoding='utf-8') as f:
        for line in f:
            original_line = line.rstrip('\n')  # 保留原始行格式(除换行符)
            stripped_line = original_line.strip()

            # 遇到分割线,重置当前标识
            if split_line_re.search(original_line):
                current_tag = None
                processed_lines.append(original_line)
                continue

            # 匹配到标识行,提取并保存标识
            tag_match = tag_line_re.search(original_line)
            if tag_match:
                current_tag = tag_match.group(1)
                processed_lines.append(original_line)
                continue

            # 普通remark行,追加当前标识
            if remark_line_re.search(original_line) and current_tag:
                # 跳过标识行和分割线的重复追加
                if not tag_line_re.search(original_line) and not split_line_re.search(original_line):
                    processed_lines.append(f"{original_line} {current_tag}")
                    continue

            # 其他行直接保留
            processed_lines.append(original_line)

    # 替换所有标识为映射词汇
    final_lines = []
    for line in processed_lines:
        replaced_line = line
        for old_tag, new_word in mapping_dict.items():
            replaced_line = re.sub(re.escape(old_tag), new_word, replaced_line)
        final_lines.append(replaced_line)

    # 写入结果文件
    with open(output_file, 'w', encoding='utf-8') as f:
        f.write('\n'.join(final_lines) + '\n')

# 调用示例
process_remark_file("file.txt", "mapping.txt", "processed_file.txt")

关键优化点

  1. 状态维护:通过current_tag变量持续跟踪当前需追加的标识,实现多段内容的循环处理
  2. 格式兼容性:保留原始行的缩进和空白,避免破坏文件原有格式
  3. 正则精准性:针对分割线、标识行、普通remark行分别编写正则,减少误匹配概率
  4. 映射效率:提前构建映射字典,批量替换时直接查询,提升处理速度

内容的提问来源于stack exchange,提问作者Adema_It_Man

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.01 20:30:38