You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python拆分含大量report的XML文件为指定数量的小XML文件?

问题描述

我有一批XML文件,每个文件中都包含数百个以XML格式存储的report,结构如下:

<maintag>
<report1>
<value1>Data</value1>
<value2>Data</value2>
<value3>Data</value3>
</report1>
<report2>
<value1>Data</value1>
<value2>Data</value2>
<value3>Data</value3>
</report2>
<report3>
<value1>Data</value1>
<value2>Data</value2>
<value3>Data</value3>
</report3>
...
</maintag>

我想要编写Python脚本将这些大文件拆分为多个小文件,每个小文件包含n个report,且这些report需包裹在<maintag>标签内。例如,若某文件包含20000个report,按每个文件100个report拆分,需得到200个结构如下的小文件:

<maintag>
<report1>
<value1>Data</value1>
<value2>Data</value2>
<value3>Data</value3>
</report1>
...
<report100>
<value1>Data</value1>
<value2>Data</value2>
<value3>Data</value3>
</report100>
</maintag>

我查阅了很多资料,大多是关于数据提取的内容,也尝试过现有的XML拆分工具,但都没有帮助。作为编程新手,我想知道这个需求是否可以用Python实现?

解决方案

完全可以用Python实现这个需求,以下是针对新手的分步实现方案:

核心思路

考虑到XML文件可能体积较大,直接用DOM解析会占用大量内存,这里采用逐行读取+字符串匹配的方式处理,既高效又节省内存,适合处理大文件。

完整代码

import os

def split_xml_large_file(input_file, reports_per_file=100):
    # 创建输出目录,避免文件混乱
    output_dir = f"{os.path.splitext(input_file)[0]}_splits"
    os.makedirs(output_dir, exist_ok=True)
    
    current_report_count = 0
    file_index = 1
    output_content = []
    in_report = False
    report_buffer = []

    with open(input_file, 'r', encoding='utf-8') as f:
        for line in f:
            stripped_line = line.strip()
            # 跳过空行
            if not stripped_line:
                continue
            
            # 匹配report开始标签(支持<report1>、<report2>等格式)
            if stripped_line.startswith('<report') and stripped_line.endswith('>'):
                in_report = True
                report_buffer.append(line)
            # 匹配report结束标签(支持</report1>、</report2>等格式)
            elif stripped_line.startswith('</report') and stripped_line.endswith('>'):
                report_buffer.append(line)
                in_report = False
                # 将完整的report加入当前输出内容
                output_content.extend(report_buffer)
                report_buffer = []
                current_report_count += 1
                
                # 达到指定数量时写入文件
                if current_report_count == reports_per_file:
                    write_split_file(output_dir, file_index, output_content)
                    # 重置计数器和内容缓存
                    current_report_count = 0
                    file_index += 1
                    output_content = []
            # 处于report内部的行,加入缓冲区
            elif in_report:
                report_buffer.append(line)
    
    # 处理剩余的不足设定数量的report
    if current_report_count > 0:
        write_split_file(output_dir, file_index, output_content)

def write_split_file(output_dir, file_index, report_content):
    # 构造带序号的输出文件名
    output_file = os.path.join(output_dir, f"split_{file_index:04d}.xml")
    # 写入标准XML结构
    with open(output_file, 'w', encoding='utf-8') as f:
        f.write("<maintag>\n")
        f.writelines(report_content)
        f.write("</maintag>\n")
    print(f"已生成文件: {output_file}")

# 使用示例
if __name__ == "__main__":
    # 替换为你的目标XML文件路径
    input_xml = "your_large_file.xml"
    # 设置每个小文件包含的report数量
    split_size = 100
    split_xml_large_file(input_xml, split_size)

代码说明

  • 内存友好:逐行读取文件,不会一次性加载整个大文件到内存,适合处理GB级别的XML文件。
  • 灵活配置:通过reports_per_file参数可自由设置每个小文件包含的report数量。
  • 自动管理目录:会在原文件同目录下生成[原文件名]_splits文件夹,统一存放拆分后的小文件。
  • 处理剩余数据:如果最后剩余的report数量不足设定值,也会生成单独的小文件,避免数据丢失。

使用步骤

  1. 将代码保存为split_xml.py文件。
  2. 修改代码中input_xml为你的目标XML文件路径,split_size为每个小文件要包含的report数量。
  3. 在命令行执行python split_xml.py即可运行脚本。

内容的提问来源于stack exchange,提问作者JOKER

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.17 09:35:15