You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何遍历根目录及其子目录批量解析XML文件并导出至CSV?

Solution for Recursively Parsing XMLs from All Subdirectories

Hey there! Great job getting the single-folder XML-to-CSV workflow working—you’re already most of the way there. The key missing piece here is replacing os.listdir() (which only scans one directory) with os.walk(), a built-in function that recursively traverses every subdirectory under your root path.

Here’s the revised code that handles all 27 subfolders, with some improvements for robustness:

import csv
import xml.etree.ElementTree as ET
import os

# Set your root directory here (the one containing all 27 subfolders)
root_path = "/Users.../your_root_directory"
csv_output_path = "/Users.../xml_extract.csv"

# Use 'with' statement to auto-close the CSV file (safer than manual close)
with open(csv_output_path, 'w', newline='') as xml_data_to_csv:
    csvwriter = csv.writer(xml_data_to_csv)
    header_written = False  # Flag to track if we've written the CSV header

    # Recursively walk through all subdirectories
    for current_dir, subdirs, files in os.walk(root_path):
        for filename in files:
            if filename.endswith('.xml'):
                full_file_path = os.path.join(current_dir, filename)
                print(f"Processing: {full_file_path}")

                # Parse the XML file
                tree = ET.parse(full_file_path)
                root = tree.getroot()

                extract_xml = []
                xml_headers = []

                # Extract data from the 'Info' elements
                for info_element in root.iter('Info'):
                    for element in info_element:
                        xml_headers.append(element.tag)
                        extract_xml.append(element.text)
                        print(f"{element.tag}: {element.text}")

                # Write header only once (first XML file processed)
                if not header_written and xml_headers:
                    csvwriter.writerow(xml_headers)
                    header_written = True

                # Write the extracted data row
                if extract_xml:
                    csvwriter.writerow(extract_xml)

print("\nAll XML files processed successfully!")

Key Changes Explained:

  • os.walk() Usage: This function iterates through every directory and subdirectory under root_path. For each iteration, it gives you:
    • current_dir: The path of the directory being scanned
    • subdirs: List of subdirectories in current_dir
    • files: List of files in current_dir
  • Automatic File Handling: Using with open(...) ensures the CSV file is closed properly even if an error occurs, avoiding data loss.
  • Header Management: The header_written flag ensures we only write the CSV column headers once (using the tags from the first XML file), instead of repeating them for every row.
  • Direct Processing: We no longer need to collect filenames in a list first—we process each XML file as we find it, which is more memory-efficient.

Quick Notes:

  • Make sure to replace root_path and csv_output_path with your actual file paths.
  • If your XML files have varying structures (different tags under Info), you might want to add logic to handle missing fields (e.g., fill with empty strings) to keep the CSV rows aligned.

内容的提问来源于stack exchange,提问作者Panda

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.30 18:09:04