Python如何跨多行读取匹配文本并提取指定字段
跨多行VPN配置字段提取实现方案
原有代码问题说明
- 逐行独立判断没有做上下文状态留存,匹配到首个关键词后无法关联后续行同属一个配置块的内容
- 匹配到关键词后直接打印整行,没有做字段值的定向截取,输出带冗余内容
- 提前把所有行转小写会破坏原始配置值的大小写格式,导致输出不符合预期
实现逻辑
针对连续块的配置文本,采用临时状态变量暂存已匹配到的字段,逐行扫描时持续补全字段,待目标字段全部提取完成后按要求格式输出,再清空临时变量进入下一个配置块的匹配,全程逐行读取不加载全量文件,适配大文件处理场景。
可直接运行的代码
# 临时字典,存储当前正在匹配的VPN配置字段 current_vpn = {} with open('/path/to/vpn.txt', 'r', encoding='utf-8') as file: for line in file: stripped_line = line.strip() # 提取VPN名称字段 if 'VPN Name:' in stripped_line: vpn_name_val = stripped_line.split('VPN Name:', 1)[1].strip() current_vpn['name'] = vpn_name_val # 已拿到VPN名称时,匹配同块的网关字段 if current_vpn.get('name') and 'Local Gateway:' in stripped_line and 'Remote Gateway:' in stripped_line: line_parts = stripped_line.split(',') local_gw_val = line_parts[0].split('Local Gateway:', 1)[1].strip() remote_gw_val = line_parts[1].split('Remote Gateway:', 1)[1].strip() current_vpn['local_gw'] = local_gw_val current_vpn['remote_gw'] = remote_gw_val # 目标字段全部收集完成,按指定格式输出 print(f"VPN Name: {current_vpn['name']}, Local Gateway: {current_vpn['local_gw']}, Remote Gateway: {current_vpn['remote_gw']}") # 清空临时存储,准备匹配下一个VPN配置块 current_vpn = {}
扩展说明
- 该方案内存占用恒定,不管文件是千行还是十万行都可以稳定运行
- 字段匹配严格绑定同一个配置块,不会出现不同VPN的名称和网关错配的问题
- 如果需要新增提取其他字段(如IKE版本、绑定接口、流量选择器名称),只需要对应增加字段截取逻辑即可
内容的提问来源于stack exchange,提问作者Daryll
相关产品推荐
相关产品推荐

