You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python类中用os.scandir分类处理文件及OOP实践优化需求

Got it, let's build a reusable, OOP-focused file categorizer using os.scandir—this will not only help you level up your Python OOP skills but also give you a practical tool you can drop into your main project. We'll structure it to handle audio (mp3), video (mp4), images, and track unsupported file extensions.

FileCategorizer Class Implementation

Let's break this down into manageable parts:

1. Class Initialization

First, we'll set up the class with a target directory, define our supported file types, and initialize storage for categorized files and unsupported extensions. Using sets for extensions makes lookups fast!

import os
from os.path import splitext

class FileCategorizer:
    def __init__(self, target_dir):
        # Validate the target directory exists
        if not os.path.isdir(target_dir):
            raise ValueError(f"Directory {target_dir} does not exist or is not a directory.")
        
        self.target_dir = target_dir
        
        # Define supported file types: key = category, value = set of extensions (lowercase)
        self.supported_types = {
            "audio": {".mp3"},
            "video": {".mp4"},
            "images": {".jpg", ".jpeg", ".png", ".gif", ".bmp"}
        }
        
        # Storage for categorized files
        self.categorized_files = {category: [] for category in self.supported_types.keys()}
        # Storage for unique unsupported extensions
        self.unsupported_extensions = set()

2. Core Scanning & Categorization Logic

We'll use os.scandir because it's more efficient than os.listdir—it fetches file metadata in bulk, which is great for large directories. We'll iterate through each entry, check if it's a file, extract its extension, and sort it into the right category (or mark it as unsupported).

def scan_directory(self):
        # Reset storage in case we scan multiple times
        self.categorized_files = {category: [] for category in self.supported_types.keys()}
        self.unsupported_extensions.clear()
        
        with os.scandir(self.target_dir) as entries:
            for entry in entries:
                if entry.is_file(follow_symlinks=False):
                    # Get file extension (lowercase to avoid case sensitivity)
                    file_ext = splitext(entry.name)[1].lower()
                    
                    # Check if extension is supported
                    categorized = False
                    for category, extensions in self.supported_types.items():
                        if file_ext in extensions:
                            self.categorized_files[category].append(entry.path)
                            categorized = True
                            break  # No need to check other categories
                    
                    # If not categorized, add extension to unsupported set
                    if not categorized:
                        # Add empty string if file has no extension
                        self.unsupported_extensions.add(file_ext if file_ext else "no_extension")

3. Helper Methods for Accessing Results

Add simple methods to retrieve the categorized files and unsupported extensions—this keeps the class's internal state encapsulated (good OOP practice!).

def get_categorized_files(self):
        """Return a copy of the categorized files dictionary to prevent external modification."""
        return {k: v.copy() for k, v in self.categorized_files.items()}
    
    def get_unsupported_extensions(self):
        """Return a sorted list of unsupported extensions."""
        return sorted(self.unsupported_extensions)

4. Integrating into Your Main Project

Here's how you'd use this class in your main code. It's straightforward and keeps your project organized:

def main():
    # Initialize the categorizer with your target directory
    categorizer = FileCategorizer("/path/to/your/files")
    
    # Run the scan
    categorizer.scan_directory()
    
    # Access results
    categorized = categorizer.get_categorized_files()
    unsupported = categorizer.get_unsupported_extensions()
    
    # Print or process the results
    print("Categorized Files:")
    for category, files in categorized.items():
        print(f"\n{category.capitalize()}:")
        for file in files:
            print(f"  - {file}")
    
    print("\nUnsupported File Extensions:")
    for ext in unsupported:
        print(f"  - {ext if ext != 'no_extension' else 'Files with no extension'}")

if __name__ == "__main__":
    main()

Bonus Tips for Improvement

  • Make it flexible: Add a method to update supported types (e.g., add_supported_type(category, extensions)) so you don't have to modify the class directly.
  • Add error handling: Wrap the os.scandir block in a try-except to handle permissions issues or other filesystem errors.
  • Recursive scanning: Extend the scan_directory method to handle subdirectories by adding a recursive parameter and using entry.is_dir() to recurse.

内容的提问来源于stack exchange,提问作者user837288

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 09:19:01