Python3:输入为文件/目录时,如何实现对应文件遍历(含XML场景)
Handling Directory and File Inputs with glob in Python
When you need to process either a directory (and all its XML files recursively) or a single XML file consistently in a for loop, the key is to first check the type of input path and generate the appropriate iterable. Here's a clean, reusable solution:
Step 1: Create a Helper Function
This function will handle both directory and file inputs, returning an iterable that works seamlessly with your for loop:
import os import glob def get_xml_files(input_path): # Check if input is a valid file if os.path.isfile(input_path): # Verify it's an XML file if input_path.lower().endswith('.xml'): yield input_path else: raise ValueError(f"Input '{input_path}' is not an XML file") # Check if input is a valid directory elif os.path.isdir(input_path): # Recursively find all XML files in the directory yield from glob.iglob(os.path.join(input_path, '**', '*.xml'), recursive=True) # Handle invalid paths else: raise FileNotFoundError(f"Input path '{input_path}' does not exist")
Step 2: Use the Function in Your For Loop
Now you can use this function to process both directories and files without modifying your loop logic:
Example 1: Input is a Directory
dirs = "foo" for file in get_xml_files(dirs): print(file)
Output:
foo/b.xml foo/c.xml foo/a.xml
Example 2: Input is an XML File
dirs = "foo/b.xml" for file in get_xml_files(dirs): print(file)
Output:
foo/b.xml
Why This Works
- Input Validation: The function first checks if the path is a file or directory, avoiding the broken glob pattern issue from your original code (where passing a file path to
{0}/**/*.xmllooks for non-existent subdirectories under the file). - Consistent Iterable: Whether you pass a directory or file, the function returns an iterator, so your for loop stays identical in both cases.
- Cross-Platform Compatibility: Using
os.path.joinensures path formatting works correctly on Windows, macOS, and Linux. - Error Handling: The function raises clear errors for invalid paths or non-XML files, making debugging easier.
Alternative: Simplified Version (Without Error Handling)
If you prefer a more concise version that skips non-XML files instead of raising errors:
import os import glob def get_xml_files(input_path): if os.path.isfile(input_path): if input_path.lower().endswith('.xml'): yield input_path elif os.path.isdir(input_path): yield from glob.iglob(os.path.join(input_path, '**', '*.xml'), recursive=True)
内容的提问来源于stack exchange,提问作者rodee
相关产品推荐
相关产品推荐

