You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何提取input.txt中<abc data="CS1">下的new与info属性值?

修改Python正则提取逻辑,仅获取指定标签下的属性值

核心思路是先精准定位到<abc data="CS1">与对应闭合标签</abc>之间的内容区块,再在这个范围内提取目标<number>标签的属性值,避免提取到其他<abc>标签下的内容。

修改后的代码示例

import gzip
import re

def extract_cs1_numbers(input_path, output_path):
    # 读取压缩文件(如果input.txt未压缩,替换为普通open即可)
    with gzip.open(input_path, 'rt', encoding='utf-8') as f:
        content = f.read()
    
    # 匹配所有<abc data="CS1">到对应</abc>的内容区块
    cs1_block_re = re.compile(r'<abc data="CS1">(.*?)</abc>', re.DOTALL)
    cs1_content_blocks = cs1_block_re.findall(content)
    
    # 定义匹配<number>标签new和info属性的正则
    number_attr_re = re.compile(r'<number new="([^"]+)" info="([^"]+)"')
    
    extracted_data = []
    for block in cs1_content_blocks:
        # 在每个CS1专属区块内提取属性值
        matches = number_attr_re.findall(block)
        extracted_data.extend(matches)
    
    # 将结果写入输出文件
    with open(output_path, 'w', encoding='utf-8') as f:
        for new_val, info_val in extracted_data:
            f.write(f"new: {new_val}, info: {info_val}\n")

# 执行提取
extract_cs1_numbers('input.txt.gz', 'output.txt')

关键修改说明

  • 先锁定目标区块:使用re.DOTALL参数让正则的.能匹配换行符,确保跨行的<abc>标签内容也能被完整捕获;非贪婪模式.*?避免误匹配到其他<abc>标签的闭合部分。
  • 精准匹配属性值:用[^"]+代替通用的.*?匹配属性值,避免因属性值内存在特殊字符(如转义引号)导致的匹配错误,提升正则的稳定性。
  • 适配非压缩文件:如果你的input.txt没有用gzip压缩,直接把gzip.open替换为普通的open即可。

内容的提问来源于stack exchange,提问作者V_S

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.12 23:40:25