You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何逐个读取并处理所有XML文件?Jupyter Notebook解析实现

Process All XML Files in a Directory with Your Existing Extraction Logic

Got it, let's turn your single-file workflow into a batch-processing system for all XMLs in your directory. Here's a practical, maintainable approach that builds on your existing code:

Step 1: Refactor Your Extraction Logic into a Reusable Function

First, wrap your data extraction code into a function that takes a single XML file path and returns the extracted data. This keeps your code clean and easy to reuse across multiple files.

from lxml import etree as ET
import os
import pandas as pd

def process_single_xml(file_path):
    # Parse the XML file
    tree = ET.parse(file_path)
    root = tree.getroot()
    
    # Extract error codes (your existing logic)
    error_codes = []
    for field in root.findall('.//Book/Message/Param/Buffer/Data/Field[11]'):
        raw_value = field.find('RawValue').text
        if raw_value is not None:
            error_codes.append(raw_value)
    
    # Add your other extraction blocks here
    # Example: Extract another field
    # other_data = []
    # for item in root.findall('.//Your/Other/XPath/Here'):
    #     value = item.find('TargetElement').text
    #     if value:
    #         other_data.append(value)
    
    # Return a dictionary with all extracted data + file name (for traceability)
    return {
        'source_file': os.path.basename(file_path),
        'error_codes': error_codes
        # 'other_data': other_data  # Uncomment if you add more fields
    }

Step 2: Fetch All XML Files in Your Directory

Use either os.listdir or glob to get paths for all .xml files in your target directory. glob is simpler because it filters files directly:

# Define your target directory
xml_directory = r'C:\Users\mysky\Documents\Decoded'

# Get all XML file paths
# Option 1: Using glob (recommended)
import glob
xml_file_paths = glob.glob(os.path.join(xml_directory, '*.xml'))

# Option 2: Using os.listdir
# xml_file_paths = [
#     os.path.join(xml_directory, filename)
#     for filename in os.listdir(xml_directory)
#     if filename.lower().endswith('.xml')
# ]

Step 3: Batch Process All Files

Loop through each XML file, run your extraction function, and collect results. Add error handling to skip corrupted files without breaking the entire workflow:

all_extracted_data = []

for file_path in xml_file_paths:
    print(f"Processing {os.path.basename(file_path)}...")
    try:
        file_results = process_single_xml(file_path)
        all_extracted_data.append(file_results)
    except Exception as e:
        print(f"⚠️ Failed to process {file_path}: {str(e)}")
        # You can log errors to a file here if needed

Step 4: Convert Results to a Pandas DataFrame

Turn the collected data into a DataFrame for easy analysis, cleaning, or export. If your extracted fields are lists (like error_codes), use explode() to expand them into individual rows:

# Create base DataFrame
df = pd.DataFrame(all_extracted_data)

# Expand list columns into separate rows (optional but useful for analysis)
df_expanded = df.explode('error_codes').reset_index(drop=True)

# Example: Export to CSV
df_expanded.to_csv('all_xml_error_codes.csv', index=False)

Key Tips

  • Path Safety: Always use os.path.join() to build file paths (avoids issues with slashes on different OSes) and prefix paths with r to escape backslashes.
  • Error Handling: The try-except block ensures a single bad XML file won't crash your entire batch job.
  • Maintainability: Keeping extraction logic in a function makes it easy to update or add new fields later.

内容的提问来源于stack exchange,提问作者M-M

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 08:19:54