Python类中用os.scandir分类处理文件及OOP实践优化需求
Got it, let's build a reusable, OOP-focused file categorizer using os.scandir—this will not only help you level up your Python OOP skills but also give you a practical tool you can drop into your main project. We'll structure it to handle audio (mp3), video (mp4), images, and track unsupported file extensions.
Let's break this down into manageable parts:
1. Class Initialization
First, we'll set up the class with a target directory, define our supported file types, and initialize storage for categorized files and unsupported extensions. Using sets for extensions makes lookups fast!
import os from os.path import splitext class FileCategorizer: def __init__(self, target_dir): # Validate the target directory exists if not os.path.isdir(target_dir): raise ValueError(f"Directory {target_dir} does not exist or is not a directory.") self.target_dir = target_dir # Define supported file types: key = category, value = set of extensions (lowercase) self.supported_types = { "audio": {".mp3"}, "video": {".mp4"}, "images": {".jpg", ".jpeg", ".png", ".gif", ".bmp"} } # Storage for categorized files self.categorized_files = {category: [] for category in self.supported_types.keys()} # Storage for unique unsupported extensions self.unsupported_extensions = set()
2. Core Scanning & Categorization Logic
We'll use os.scandir because it's more efficient than os.listdir—it fetches file metadata in bulk, which is great for large directories. We'll iterate through each entry, check if it's a file, extract its extension, and sort it into the right category (or mark it as unsupported).
def scan_directory(self): # Reset storage in case we scan multiple times self.categorized_files = {category: [] for category in self.supported_types.keys()} self.unsupported_extensions.clear() with os.scandir(self.target_dir) as entries: for entry in entries: if entry.is_file(follow_symlinks=False): # Get file extension (lowercase to avoid case sensitivity) file_ext = splitext(entry.name)[1].lower() # Check if extension is supported categorized = False for category, extensions in self.supported_types.items(): if file_ext in extensions: self.categorized_files[category].append(entry.path) categorized = True break # No need to check other categories # If not categorized, add extension to unsupported set if not categorized: # Add empty string if file has no extension self.unsupported_extensions.add(file_ext if file_ext else "no_extension")
3. Helper Methods for Accessing Results
Add simple methods to retrieve the categorized files and unsupported extensions—this keeps the class's internal state encapsulated (good OOP practice!).
def get_categorized_files(self): """Return a copy of the categorized files dictionary to prevent external modification.""" return {k: v.copy() for k, v in self.categorized_files.items()} def get_unsupported_extensions(self): """Return a sorted list of unsupported extensions.""" return sorted(self.unsupported_extensions)
4. Integrating into Your Main Project
Here's how you'd use this class in your main code. It's straightforward and keeps your project organized:
def main(): # Initialize the categorizer with your target directory categorizer = FileCategorizer("/path/to/your/files") # Run the scan categorizer.scan_directory() # Access results categorized = categorizer.get_categorized_files() unsupported = categorizer.get_unsupported_extensions() # Print or process the results print("Categorized Files:") for category, files in categorized.items(): print(f"\n{category.capitalize()}:") for file in files: print(f" - {file}") print("\nUnsupported File Extensions:") for ext in unsupported: print(f" - {ext if ext != 'no_extension' else 'Files with no extension'}") if __name__ == "__main__": main()
Bonus Tips for Improvement
- Make it flexible: Add a method to update supported types (e.g.,
add_supported_type(category, extensions)) so you don't have to modify the class directly. - Add error handling: Wrap the
os.scandirblock in a try-except to handle permissions issues or other filesystem errors. - Recursive scanning: Extend the
scan_directorymethod to handle subdirectories by adding arecursiveparameter and usingentry.is_dir()to recurse.
内容的提问来源于stack exchange,提问作者user837288

