如何批量处理目录中所有XML文件以去除空白行?
Solution to Process All XML Files and Remove Blank Lines
Got it, let's adjust your code to handle every XML file in your target directory smoothly! Here's a step-by-step breakdown and the revised working code:
Step 1: Fix Basics & Define Paths Cleanly
First, make sure you import the os module (your original code uses it but doesn’t include the import). We’ll also separate input/output directories into variables for easier maintenance, and ensure the output directory exists (so you don’t get errors if it’s missing).
Step 2: Loop Through All Target Files
Instead of hardcoding a single filename, we’ll loop through every file in your input directory. We’ll add a check to only process .xml files too—this avoids accidentally handling any non-XML files that might end up in the folder.
Revised Full Code
import os # Define your input and output directories (easy to modify later) input_dir = r'C:\Users\Max12\Desktop\xml\pdfminer\UiPath\output' output_dir = r'C:\Users\Max12\Desktop\xml\pdfminer\UiPath\out' # Create output directory if it doesn't exist (no manual setup needed!) os.makedirs(output_dir, exist_ok=True) # Loop through each file in the input directory for filename in os.listdir(input_dir): # Only process XML files to avoid unwanted files if filename.endswith('.xml'): # Build full paths for input and output files input_file_path = os.path.join(input_dir, filename) output_file_path = os.path.join(output_dir, filename) # Apply your blank-line-removal logic to each file with open(input_file_path, 'r') as infile, open(output_file_path, 'w') as outfile: for line in infile: # Skip empty lines (after stripping whitespace) if not line.strip(): continue outfile.write(line) print("All XML files processed successfully!")
What This Does
- Imports
os: Required for directory and file path operations. - Auto-creates Output Folder:
os.makedirs(..., exist_ok=True)ensures the output directory exists—no need to manually create it beforehand. - Preserves Original Filenames: Each cleaned file keeps its original name in the output folder (so
0.xmlbecomesout/0.xml,1.xmlbecomesout/1.xml, etc.), which matches your goal of 4 cleaned XML files. - Safely Filters Files: The
.endswith('.xml')check ensures we only process the files you care about. - Reuses Your Working Logic: The blank-line-removal code you already wrote is applied to every file individually.
内容的提问来源于stack exchange,提问作者Max FH
相关产品推荐
相关产品推荐

