如何用Python脚本检测文件夹中空Excel、PDF、CSV文件并导出文件名
Batch Detect Empty Excel, PDF, and CSV Files & Export to TXT
I get it—relying on file size alone doesn’t cut it here because different file types have minimal sizes even when "empty" (like those 8KB Excel files that look empty but have a default structure). Let’s build a Python script that properly checks each file type for actual content, then exports the list of empty files to a text file.
First, we’ll need a couple of libraries to handle Excel and PDF files—install them with:
pip install openpyxl PyPDF2
Here’s the complete script, with comments explaining each part:
import os import openpyxl from PyPDF2 import PdfReader import csv def is_empty_excel(file_path): """Check if Excel file has no actual data (all sheets are empty)""" try: workbook = openpyxl.load_workbook(file_path, read_only=True) for sheet_name in workbook.sheetnames: sheet = workbook[sheet_name] # Check every row for non-empty content for row in sheet.iter_rows(values_only=True): if any(cell is not None and str(cell).strip() != "" for cell in row): workbook.close() return False workbook.close() return True except Exception as e: print(f"Error reading Excel file {file_path}: {str(e)}") return False # Mark corrupted files as non-empty (adjust if needed) def is_empty_pdf(file_path): """Check if PDF has no text content or is zero bytes""" # First check for zero-byte files if os.stat(file_path).st_size == 0: return True try: reader = PdfReader(file_path) # Check if there are no pages if len(reader.pages) == 0: return True # Check if all pages have no extractable text total_text = "" for page in reader.pages: total_text += page.extract_text() or "" return len(total_text.strip()) == 0 except Exception as e: print(f"Error reading PDF file {file_path}: {str(e)}") return False def is_empty_csv(file_path, consider_header_only=False): """Check if CSV is empty, or header-only (if enabled)""" if os.stat(file_path).st_size == 0: return True try: with open(file_path, 'r', newline='', encoding='utf-8') as f: reader = csv.reader(f) rows = list(reader) # Count header-only files as empty if requested if consider_header_only: return len(rows) <= 1 # Otherwise, check for zero rows return len(rows) == 0 except Exception as e: print(f"Error reading CSV file {file_path}: {str(e)}") return False def scan_folder(folder_path): """Scan target folder and collect paths to empty files""" empty_files = [] for root, _, files in os.walk(folder_path): for file in files: file_path = os.path.join(root, file) lower_filename = file.lower() if lower_filename.endswith(('.xlsx', '.xls')): if is_empty_excel(file_path): empty_files.append(file_path) elif lower_filename.endswith('.pdf'): if is_empty_pdf(file_path): empty_files.append(file_path) elif lower_filename.endswith('.csv'): # Set consider_header_only=True if you want to flag header-only CSVs if is_empty_csv(file_path, consider_header_only=False): empty_files.append(file_path) return empty_files if __name__ == "__main__": # Replace this with your actual folder path target_folder = "./your_downloaded_files" empty_files_list = scan_folder(target_folder) # Export results to a text file output_file = "empty_files_report.txt" with open(output_file, 'w', encoding='utf-8') as f: for file_path in empty_files_list: f.write(f"{file_path}\n") print(f"Scan complete! Found {len(empty_files_list)} empty files. Report saved to {output_file}")
Key Details:
- Excel Files: Checks every sheet for non-empty cells, so it correctly identifies those 8KB "empty" files that have no actual data.
- PDF Files: Catches both zero-byte PDFs and blank-page PDFs with no extractable text.
- CSV Files: You can toggle the
consider_header_onlyparameter to decide if files with only headers count as empty. - Error Handling: Skips corrupted files and prints a message instead of crashing the script.
How to Use:
- Install the required libraries with the pip command above.
- Replace
"./your_downloaded_files"with the path to your folder of PDFs, CSVs, and Excel files. - Run the script—it will generate
empty_files_report.txtwith all paths to empty files.
This should save you the tedious work of manually checking each file!
内容的提问来源于stack exchange,提问作者Mitesh Agrawal
相关产品推荐
相关产品推荐

