You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何修改XML拆分脚本批量处理指定前缀文件并自定义输出命名

Revised XML Splitting Script for Multiple Files

Got it, let's tweak your script to handle all BGSM_VPAY-prefixed files in the folder and name outputs with the original base name plus sequential numbers. Here's the updated code, followed by breakdowns of the key changes:

import xml.etree.ElementTree as ET
import os
from glob import glob

# Grab all XML files starting with BGSM_VPAY in the current directory
for input_file in glob("BGSM_VPAY*.xml"):
    # Split original filename into base name and extension (e.g., "BGSM_VPAY_xxx" and ".xml")
    base_name, ext = os.path.splitext(input_file)
    index = 0
    
    # Fix the incomplete events parameter from your original script
    context = ET.iterparse(input_file, events=('end',))
    for event, elem in context:
        if elem.tag == 'Payment_Ack':
            index += 1
            # Generate output filename with zero-padded sequence number
            output_filename = f"{base_name}_{index:03d}{ext}"
            with open(output_filename, 'wb') as f:
                # Write proper XML declaration and the element content
                f.write(b'<?xml version="1.0" encoding="UTF-8"?>\n')
                f.write(ET.tostring(elem))
            # Clear processed element to save memory (critical for large files)
            elem.clear()

Key Changes Explained:

  • Batch file handling: Used glob("BGSM_VPAY*.xml") to automatically detect all matching XML files in the current folder, so you don't have to hardcode filenames one by one.
  • Smart output naming:
    • os.path.splitext(input_file) splits the original filename into its base (no extension) and extension parts.
    • f"{base_name}_{index:03d}{ext}" creates filenames like BGSM_VPAY_D-001565_20180315-220009-049_001.xml — the :03d ensures sequence numbers are zero-padded to 3 digits (001, 002, etc.) for clean sorting.
  • Syntax fix: Corrected the broken events parameter in ET.iterparse (your original had an incomplete string 'e$) which would throw an error).
  • Memory optimization: Added elem.clear() after writing each element — this prevents the script from hogging RAM when processing large XML files by clearing processed elements from memory.

内容的提问来源于stack exchange,提问作者dhoggard

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 03:49:28