You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python脚本检测文件夹中空Excel、PDF、CSV文件并导出文件名

Batch Detect Empty Excel, PDF, and CSV Files & Export to TXT

I get it—relying on file size alone doesn’t cut it here because different file types have minimal sizes even when "empty" (like those 8KB Excel files that look empty but have a default structure). Let’s build a Python script that properly checks each file type for actual content, then exports the list of empty files to a text file.

First, we’ll need a couple of libraries to handle Excel and PDF files—install them with:

pip install openpyxl PyPDF2

Here’s the complete script, with comments explaining each part:

import os
import openpyxl
from PyPDF2 import PdfReader
import csv

def is_empty_excel(file_path):
    """Check if Excel file has no actual data (all sheets are empty)"""
    try:
        workbook = openpyxl.load_workbook(file_path, read_only=True)
        for sheet_name in workbook.sheetnames:
            sheet = workbook[sheet_name]
            # Check every row for non-empty content
            for row in sheet.iter_rows(values_only=True):
                if any(cell is not None and str(cell).strip() != "" for cell in row):
                    workbook.close()
                    return False
        workbook.close()
        return True
    except Exception as e:
        print(f"Error reading Excel file {file_path}: {str(e)}")
        return False  # Mark corrupted files as non-empty (adjust if needed)

def is_empty_pdf(file_path):
    """Check if PDF has no text content or is zero bytes"""
    # First check for zero-byte files
    if os.stat(file_path).st_size == 0:
        return True
    try:
        reader = PdfReader(file_path)
        # Check if there are no pages
        if len(reader.pages) == 0:
            return True
        # Check if all pages have no extractable text
        total_text = ""
        for page in reader.pages:
            total_text += page.extract_text() or ""
        return len(total_text.strip()) == 0
    except Exception as e:
        print(f"Error reading PDF file {file_path}: {str(e)}")
        return False

def is_empty_csv(file_path, consider_header_only=False):
    """Check if CSV is empty, or header-only (if enabled)"""
    if os.stat(file_path).st_size == 0:
        return True
    try:
        with open(file_path, 'r', newline='', encoding='utf-8') as f:
            reader = csv.reader(f)
            rows = list(reader)
            # Count header-only files as empty if requested
            if consider_header_only:
                return len(rows) <= 1
            # Otherwise, check for zero rows
            return len(rows) == 0
    except Exception as e:
        print(f"Error reading CSV file {file_path}: {str(e)}")
        return False

def scan_folder(folder_path):
    """Scan target folder and collect paths to empty files"""
    empty_files = []
    for root, _, files in os.walk(folder_path):
        for file in files:
            file_path = os.path.join(root, file)
            lower_filename = file.lower()
            
            if lower_filename.endswith(('.xlsx', '.xls')):
                if is_empty_excel(file_path):
                    empty_files.append(file_path)
            elif lower_filename.endswith('.pdf'):
                if is_empty_pdf(file_path):
                    empty_files.append(file_path)
            elif lower_filename.endswith('.csv'):
                # Set consider_header_only=True if you want to flag header-only CSVs
                if is_empty_csv(file_path, consider_header_only=False):
                    empty_files.append(file_path)
    return empty_files

if __name__ == "__main__":
    # Replace this with your actual folder path
    target_folder = "./your_downloaded_files"
    empty_files_list = scan_folder(target_folder)
    
    # Export results to a text file
    output_file = "empty_files_report.txt"
    with open(output_file, 'w', encoding='utf-8') as f:
        for file_path in empty_files_list:
            f.write(f"{file_path}\n")
    
    print(f"Scan complete! Found {len(empty_files_list)} empty files. Report saved to {output_file}")

Key Details:

  • Excel Files: Checks every sheet for non-empty cells, so it correctly identifies those 8KB "empty" files that have no actual data.
  • PDF Files: Catches both zero-byte PDFs and blank-page PDFs with no extractable text.
  • CSV Files: You can toggle the consider_header_only parameter to decide if files with only headers count as empty.
  • Error Handling: Skips corrupted files and prints a message instead of crashing the script.

How to Use:

  1. Install the required libraries with the pip command above.
  2. Replace "./your_downloaded_files" with the path to your folder of PDFs, CSVs, and Excel files.
  3. Run the script—it will generate empty_files_report.txt with all paths to empty files.

This should save you the tedious work of manually checking each file!

内容的提问来源于stack exchange,提问作者Mitesh Agrawal

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.13 09:28:43