You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从gzip压缩文件提取指定属性值并写入目标文件?

问题需求

从input.txt.gz压缩文件的类XML文本行中,提取<number>标签内的new、info、no_id属性值,按指定格式写入output.txt文件,但现有Python代码无法完成该任务。

输入文件示例

<hello script="2.5">
<welcome>
     <hgsdhjaghjdghjagdjhgjdhgdajhgdajhgdhjjgfkjg
     <number new="0x0000-0x3FF" Id="bhi" Range="4" no_id="CS_hello" />
               
          <----jsdjhsdjndkjjdhjdJHksdkjdnknnddnekfgrejgjorgj jregjgkrjglrjgojggjorjg---&gt;
          <number new="0x02" Id="bhi" Unit="0" Range="4" info="0x00000012" no_id="hi_all res" />
          <number new="0x04" Id="bhi" Unit="0" Range="4" info="0x0000023f" no_id="d. hwd mas" />
          <---- dfiuhdwiudi iwqdidffenfj odwqjdjqwgru jdqkkjwfkjfwn odHHOIJD JSDNKS nsk----&gt;
          <number new="0x06" Id="bhi" Unit="0" Range="4" info="0x00000f22" no_id="sjkdnkl jdsnj (Sedk)" />
          <number new="0x08" Id="bhi" Unit="0" Range="4" info="0x00000f1b" no_id="dm o_1_3k_2_0" />
    <---bdheuh jwdhjwdkiwh----&gt;
          <number new="0x32" Id="bhi"  Range="4" info="0x000012f5" no_id="HES kd" />
          <number new="0x336" Id="bhi" Range="4" info="0x00000df2" no_id="dnkwn" />
<--adhhj jdwjdkkj jsSDjkasdj jefnflefk kjsjfoekfle kajfofkp ksaokdfpef----&gt;
<---the end of file----&gt;

期望输出示例

new="0x02" info="0x00000012" no_id="hi_all res"
new="0x04" info="0x0000023f" no_id="d. hwd mas"
new="0x06" info="0x00000f22" no_id="sjkdnkl jdsnj (Sedk)"
new="0x08" info="0x00000f1b" no_id="dm o_1_3k_2_0"
new="0x32" info="0x000012f5" no_id="HES kd"
new="0x336" info="0x00000df2" no_id="dnkwn"

当前代码问题

现有代码通过分割空格提取属性的方式存在多处缺陷:

  • 属性在标签中的顺序不固定,依赖索引(如cols[1]、cols[5])会导致提取错误
  • 属性值可能包含空格(如no_id="hi_all res"),直接分割空格会破坏属性值的完整性
  • 未提取info属性,也未将结果写入output.txt,仅做打印输出
  • 字符串处理逻辑存在错误(strip参数格式无效)
import gzip
with gzip.open("input.txt.gz", "rb") as fin:
     with open("output.txt", "w") as fout:
           for line in fin:
                if line.decode('utf-8').strip():
                   line = line.decode('utf-8').strip("\n' '")
                   cols = line.split(" ")
                   if len(cols) >= 5:
                      print(cols[1], cols[5])

解决方案代码

使用正则表达式匹配属性,能稳定处理属性顺序变化、属性值含空格的场景:

import gzip
import re

# 定义匹配三个属性的正则模式,兼容属性顺序变化
pattern = re.compile(r'new="([^"]+)"\s.*?info="([^"]+)"\s.*?no_id="([^"]+)"')

with gzip.open("input.txt.gz", "rb") as fin, open("output.txt", "w", encoding="utf-8") as fout:
    for line_bytes in fin:
        line = line_bytes.decode("utf-8").strip()
        # 只处理包含<number>且带有info属性的行
        if "<number" in line and "info=" in line:
            match = pattern.search(line)
            if match:
                new_val, info_val, no_id_val = match.groups()
                # 按指定格式拼接并写入文件
                output_line = f'new="{new_val}" info="{info_val}" no_id="{no_id_val}"\n'
                fout.write(output_line)

代码说明

  1. 正则模式精准匹配三个属性,不受它们在标签中的顺序影响
  2. 过滤掉无info属性的<number>行,符合输出示例要求
  3. 正确解码压缩文件内容,直接将结果写入目标文件
  4. 支持属性值包含空格、特殊字符的场景

内容的提问来源于stack exchange,提问作者V_S

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.15 08:20:53