如何逐个读取并处理所有XML文件?Jupyter Notebook解析实现
Got it, let's turn your single-file workflow into a batch-processing system for all XMLs in your directory. Here's a practical, maintainable approach that builds on your existing code:
Step 1: Refactor Your Extraction Logic into a Reusable Function
First, wrap your data extraction code into a function that takes a single XML file path and returns the extracted data. This keeps your code clean and easy to reuse across multiple files.
from lxml import etree as ET import os import pandas as pd def process_single_xml(file_path): # Parse the XML file tree = ET.parse(file_path) root = tree.getroot() # Extract error codes (your existing logic) error_codes = [] for field in root.findall('.//Book/Message/Param/Buffer/Data/Field[11]'): raw_value = field.find('RawValue').text if raw_value is not None: error_codes.append(raw_value) # Add your other extraction blocks here # Example: Extract another field # other_data = [] # for item in root.findall('.//Your/Other/XPath/Here'): # value = item.find('TargetElement').text # if value: # other_data.append(value) # Return a dictionary with all extracted data + file name (for traceability) return { 'source_file': os.path.basename(file_path), 'error_codes': error_codes # 'other_data': other_data # Uncomment if you add more fields }
Step 2: Fetch All XML Files in Your Directory
Use either os.listdir or glob to get paths for all .xml files in your target directory. glob is simpler because it filters files directly:
# Define your target directory xml_directory = r'C:\Users\mysky\Documents\Decoded' # Get all XML file paths # Option 1: Using glob (recommended) import glob xml_file_paths = glob.glob(os.path.join(xml_directory, '*.xml')) # Option 2: Using os.listdir # xml_file_paths = [ # os.path.join(xml_directory, filename) # for filename in os.listdir(xml_directory) # if filename.lower().endswith('.xml') # ]
Step 3: Batch Process All Files
Loop through each XML file, run your extraction function, and collect results. Add error handling to skip corrupted files without breaking the entire workflow:
all_extracted_data = [] for file_path in xml_file_paths: print(f"Processing {os.path.basename(file_path)}...") try: file_results = process_single_xml(file_path) all_extracted_data.append(file_results) except Exception as e: print(f"⚠️ Failed to process {file_path}: {str(e)}") # You can log errors to a file here if needed
Step 4: Convert Results to a Pandas DataFrame
Turn the collected data into a DataFrame for easy analysis, cleaning, or export. If your extracted fields are lists (like error_codes), use explode() to expand them into individual rows:
# Create base DataFrame df = pd.DataFrame(all_extracted_data) # Expand list columns into separate rows (optional but useful for analysis) df_expanded = df.explode('error_codes').reset_index(drop=True) # Example: Export to CSV df_expanded.to_csv('all_xml_error_codes.csv', index=False)
Key Tips
- Path Safety: Always use
os.path.join()to build file paths (avoids issues with slashes on different OSes) and prefix paths withrto escape backslashes. - Error Handling: The
try-exceptblock ensures a single bad XML file won't crash your entire batch job. - Maintainability: Keeping extraction logic in a function makes it easy to update or add new fields later.
内容的提问来源于stack exchange,提问作者M-M

