如何提取input.txt中<abc data="CS1">下的new与info属性值?
修改Python正则提取逻辑,仅获取指定标签下的属性值
核心思路是先精准定位到<abc data="CS1">与对应闭合标签</abc>之间的内容区块,再在这个范围内提取目标<number>标签的属性值,避免提取到其他<abc>标签下的内容。
修改后的代码示例
import gzip import re def extract_cs1_numbers(input_path, output_path): # 读取压缩文件(如果input.txt未压缩,替换为普通open即可) with gzip.open(input_path, 'rt', encoding='utf-8') as f: content = f.read() # 匹配所有<abc data="CS1">到对应</abc>的内容区块 cs1_block_re = re.compile(r'<abc data="CS1">(.*?)</abc>', re.DOTALL) cs1_content_blocks = cs1_block_re.findall(content) # 定义匹配<number>标签new和info属性的正则 number_attr_re = re.compile(r'<number new="([^"]+)" info="([^"]+)"') extracted_data = [] for block in cs1_content_blocks: # 在每个CS1专属区块内提取属性值 matches = number_attr_re.findall(block) extracted_data.extend(matches) # 将结果写入输出文件 with open(output_path, 'w', encoding='utf-8') as f: for new_val, info_val in extracted_data: f.write(f"new: {new_val}, info: {info_val}\n") # 执行提取 extract_cs1_numbers('input.txt.gz', 'output.txt')
关键修改说明
- 先锁定目标区块:使用
re.DOTALL参数让正则的.能匹配换行符,确保跨行的<abc>标签内容也能被完整捕获;非贪婪模式.*?避免误匹配到其他<abc>标签的闭合部分。 - 精准匹配属性值:用
[^"]+代替通用的.*?匹配属性值,避免因属性值内存在特殊字符(如转义引号)导致的匹配错误,提升正则的稳定性。 - 适配非压缩文件:如果你的
input.txt没有用gzip压缩,直接把gzip.open替换为普通的open即可。
内容的提问来源于stack exchange,提问作者V_S
相关产品推荐
相关产品推荐

