如何遍历根目录及其子目录批量解析XML文件并导出至CSV?
Solution for Recursively Parsing XMLs from All Subdirectories
Hey there! Great job getting the single-folder XML-to-CSV workflow working—you’re already most of the way there. The key missing piece here is replacing os.listdir() (which only scans one directory) with os.walk(), a built-in function that recursively traverses every subdirectory under your root path.
Here’s the revised code that handles all 27 subfolders, with some improvements for robustness:
import csv import xml.etree.ElementTree as ET import os # Set your root directory here (the one containing all 27 subfolders) root_path = "/Users.../your_root_directory" csv_output_path = "/Users.../xml_extract.csv" # Use 'with' statement to auto-close the CSV file (safer than manual close) with open(csv_output_path, 'w', newline='') as xml_data_to_csv: csvwriter = csv.writer(xml_data_to_csv) header_written = False # Flag to track if we've written the CSV header # Recursively walk through all subdirectories for current_dir, subdirs, files in os.walk(root_path): for filename in files: if filename.endswith('.xml'): full_file_path = os.path.join(current_dir, filename) print(f"Processing: {full_file_path}") # Parse the XML file tree = ET.parse(full_file_path) root = tree.getroot() extract_xml = [] xml_headers = [] # Extract data from the 'Info' elements for info_element in root.iter('Info'): for element in info_element: xml_headers.append(element.tag) extract_xml.append(element.text) print(f"{element.tag}: {element.text}") # Write header only once (first XML file processed) if not header_written and xml_headers: csvwriter.writerow(xml_headers) header_written = True # Write the extracted data row if extract_xml: csvwriter.writerow(extract_xml) print("\nAll XML files processed successfully!")
Key Changes Explained:
os.walk()Usage: This function iterates through every directory and subdirectory underroot_path. For each iteration, it gives you:current_dir: The path of the directory being scannedsubdirs: List of subdirectories incurrent_dirfiles: List of files incurrent_dir
- Automatic File Handling: Using
with open(...)ensures the CSV file is closed properly even if an error occurs, avoiding data loss. - Header Management: The
header_writtenflag ensures we only write the CSV column headers once (using the tags from the first XML file), instead of repeating them for every row. - Direct Processing: We no longer need to collect filenames in a list first—we process each XML file as we find it, which is more memory-efficient.
Quick Notes:
- Make sure to replace
root_pathandcsv_output_pathwith your actual file paths. - If your XML files have varying structures (different tags under
Info), you might want to add logic to handle missing fields (e.g., fill with empty strings) to keep the CSV rows aligned.
内容的提问来源于stack exchange,提问作者Panda
相关产品推荐
相关产品推荐

