如何用Python批量拆分目录中所有多页PDF为单页文件?
Batch Split All PDFs in a Directory with Python
Great job getting the single-file PDF splitter working! Extending this to handle all PDFs in a directory is straightforward—we just need to add a loop to iterate through every PDF file and run your existing pdf_splitter function on each one.
Here's the updated code that does exactly that:
# pdf_splitter.py import os from PyPDF2 import PdfFileReader, PdfFileWriter def pdf_splitter(path): fname = os.path.splitext(os.path.basename(path))[0] # Optional: Create a subdirectory to store split pages (keeps things organized) output_dir = f"{fname}_pages" os.makedirs(output_dir, exist_ok=True) pdf = PdfFileReader(path) for page in range(pdf.getNumPages()): pdf_writer = PdfFileWriter() pdf_writer.addPage(pdf.getPage(page)) # Save pages to the dedicated subdirectory output_filename = os.path.join(output_dir, f"{fname}_page_{page+1}.pdf") with open(output_filename, 'wb') as out: pdf_writer.write(out) print(f'Created: {output_filename}') def batch_split_pdfs(directory): # Loop through all files in the target directory for filename in os.listdir(directory): # Only process PDF files (case-insensitive check) if filename.lower().endswith('.pdf'): # Skip already split pages (optional, remove if you want to reprocess all PDFs) if '_page_' not in filename: file_path = os.path.join(directory, filename) print(f'\nProcessing {filename}...') pdf_splitter(file_path) if __name__ == '__main__': # Use the current working directory by default # Replace with your target path if needed, e.g., 'C:/Documents/MyPDFs' target_directory = os.getcwd() batch_split_pdfs(target_directory)
Key Improvements & Notes:
- Batch Processing: The
batch_split_pdfsfunction scans your target directory and processes every PDF file it finds. - Case Insensitivity: Checks for
.pdfor.PDFextensions so no files are missed. - Organized Output: Added a subdirectory for each original PDF's split pages (e.g.,
w9.pdfgets aw9_pagesfolder with all its single-page files). Remove this part if you prefer all files in the root directory. - Avoid Reprocessing: Skips files that already have
_page_in their name (so you don't split the split pages by accident). Remove that check if you want to process every PDF regardless.
To use this:
- Save the updated code as
pdf_splitter.py - Place it in the directory with your PDFs (or update
target_directoryto point to your PDF folder) - Run the script—all your PDFs will be split into single-page files automatically!
内容的提问来源于stack exchange,提问作者Bob Maiers
相关产品推荐
相关产品推荐

