You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python批量未知文件处理的测试方案及实例求助

Got it, let's build a practical Python workflow that handles unknown files with built-in testing, error detection, and a final report—starting with an initial sanity check before any processing begins. This example is flexible enough to adapt to your specific file types and processing needs.

Core Workflow

The system follows these steps:

  1. Initialization Test: Validate the processing environment (e.g., directory access, empty input folder checks)
  2. File Discovery: Scan the target directory for all unknown files
  3. Per-File Testing: Run validation checks on each file before processing
  4. Safe Processing: Attempt to process files that pass tests, catching errors if they occur
  5. Report Generation: Compile a detailed report of successes, failures, and error reasons
Full Python Implementation

Here's a complete, runnable example. You'll need to install python-magic first for file type detection (run pip install python-magic-bin on Windows, or pip install python-magic on macOS/Linux):

import os
import json
import hashlib
import magic
from datetime import datetime

def initialization_test(input_dir, output_dir):
    """Run initial checks to ensure the environment is ready"""
    test_results = {
        "timestamp": datetime.now().isoformat(),
        "passed": True,
        "errors": []
    }

    # Check if input directory exists
    if not os.path.exists(input_dir):
        test_results["passed"] = False
        test_results["errors"].append(f"Input directory {input_dir} does not exist")
    else:
        # Check if input directory is readable
        if not os.access(input_dir, os.R_OK):
            test_results["passed"] = False
            test_results["errors"].append(f"No read access to input directory {input_dir}")
        # Check if input directory is empty (optional, adjust based on your needs)
        if len(os.listdir(input_dir)) == 0:
            test_results["errors"].append(f"Input directory {input_dir} is empty")

    # Check if output directory exists, create if not
    if not os.path.exists(output_dir):
        try:
            os.makedirs(output_dir)
            test_results["errors"].append(f"Created output directory {output_dir}")
        except OSError as e:
            test_results["passed"] = False
            test_results["errors"].append(f"Failed to create output directory: {str(e)}")
    else:
        if not os.access(output_dir, os.W_OK):
            test_results["passed"] = False
            test_results["errors"].append(f"No write access to output directory {output_dir}")

    return test_results

def validate_file(file_path):
    """Run validation tests on a single unknown file"""
    validation = {
        "file_path": file_path,
        "valid": True,
        "issues": []
    }

    # Test 1: Check if file is readable
    if not os.access(file_path, os.R_OK):
        validation["valid"] = False
        validation["issues"].append("Cannot read file (permission denied)")
        return validation

    # Test 2: Check file size (skip zero-byte files)
    if os.path.getsize(file_path) == 0:
        validation["valid"] = False
        validation["issues"].append("Zero-byte file, skipping")
        return validation

    # Test 3: Detect file type (adjust based on your expected types)
    try:
        file_type = magic.from_file(file_path, mime=True)
        # Example: Allow only text or image files (modify this to your needs)
        allowed_types = ["text/plain", "image/jpeg", "image/png"]
        if file_type not in allowed_types:
            validation["issues"].append(f"Unsupported file type: {file_type}")
            # Optional: Mark as invalid if you want to skip non-allowed types
            # validation["valid"] = False
    except Exception as e:
        validation["issues"].append(f"Failed to detect file type: {str(e)}")

    # Test 4: Check file integrity (optional, using MD5 hash example)
    try:
        with open(file_path, "rb") as f:
            hash_obj = hashlib.md5()
            while chunk := f.read(4096):
                hash_obj.update(chunk)
            validation["md5_hash"] = hash_obj.hexdigest()
    except Exception as e:
        validation["issues"].append(f"Failed to compute file hash: {str(e)}")

    return validation

def process_file(file_path, output_dir):
    """Example processing function - replace with your actual logic"""
    try:
        # Example: Copy valid files to output directory (replace with your processing)
        file_name = os.path.basename(file_path)
        output_path = os.path.join(output_dir, file_name)
        with open(file_path, "rb") as src, open(output_path, "wb") as dst:
            dst.write(src.read())
        return {"success": True, "message": f"Successfully copied to {output_path}"}
    except Exception as e:
        return {"success": False, "error": str(e)}

def generate_final_report(init_test, file_results):
    """Compile all results into a structured report"""
    report = {
        "report_timestamp": datetime.now().isoformat(),
        "initialization_test": init_test,
        "total_files_scanned": len(file_results),
        "files_processed_successfully": sum(1 for res in file_results if res["processing"].get("success")),
        "files_failed_validation": sum(1 for res in file_results if not res["validation"]["valid"]),
        "file_details": file_results
    }

    # Save report to JSON file
    report_path = os.path.join(os.getcwd(), "processing_report.json")
    with open(report_path, "w") as f:
        json.dump(report, f, indent=2)

    print(f"Final report saved to {report_path}")
    return report

def main(input_dir="./input_files", output_dir="./processed_files"):
    # Step 1: Run initialization test
    print("Running initialization tests...")
    init_test = initialization_test(input_dir, output_dir)
    if not init_test["passed"]:
        print("Initialization tests failed. Aborting process.")
        print("Initialization errors:", init_test["errors"])
        return

    # Step 2: Discover all files in input directory
    print("\nScanning input directory for files...")
    file_paths = [os.path.join(input_dir, f) for f in os.listdir(input_dir) if os.path.isfile(os.path.join(input_dir, f))]
    if not file_paths:
        print("No files found in input directory.")
        return

    # Step 3: Process each file with validation
    file_results = []
    for file_path in file_paths:
        print(f"\nProcessing {file_path}...")
        # Run validation tests
        validation = validate_file(file_path)
        if not validation["valid"]:
            print(f"Validation failed for {file_path}: {validation['issues']}")
            file_results.append({"validation": validation, "processing": {"success": False, "reason": "Validation failed"}})
            continue

        # If validation passes, attempt processing
        processing_result = process_file(file_path, output_dir)
        file_results.append({"validation": validation, "processing": processing_result})
        if processing_result["success"]:
            print(f"Processing succeeded: {processing_result['message']}")
        else:
            print(f"Processing failed: {processing_result['error']}")

    # Step 4: Generate final report
    print("\nGenerating final report...")
    generate_final_report(init_test, file_results)

if __name__ == "__main__":
    main()
Key Components Explained

Let's break down the critical parts so you can adapt them to your use case:

  • Initialization Test: This ensures your environment is stable before touching any files. You can add more checks here (e.g., checking for required dependencies, database connections) based on your workflow.
  • File Validation: The validate_file function is where you'll add custom tests for your unknown files. For example, if you're expecting CSV files, add a check for valid CSV formatting; if handling archives, verify they're not corrupted.
  • Safe Processing: The process_file function wraps your actual logic in error handling to catch unexpected issues without crashing the entire workflow. Replace the copy logic with your own processing (e.g., parsing, transforming, uploading).
  • Reporting: The final JSON report gives you a complete audit trail—you can modify it to generate HTML reports, send emails, or integrate with monitoring tools.
How to Customize This
  1. Adjust Validation Rules: Modify the validate_file function to include tests specific to your file types (e.g., XML schema validation, JSON syntax checks).
  2. Update Processing Logic: Replace the file copy in process_file with your actual file handling code (e.g., using pandas to parse CSVs, PIL to edit images).
  3. Modify Report Format: Change the JSON output to a text file, HTML, or even a database entry based on how you want to review results.
  4. Add Notifications: Extend the generate_final_report function to send an email alert if a certain number of failures occur.

内容的提问来源于stack exchange,提问作者Cirrith

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 07:05:19