Python实现PDF页面范围拆分:寻求替代pdftk+subprocess的原生库方案
Got it, let's ditch those pdftk subprocess calls and use native Python libraries to extract specific page ranges from your PDF. Two solid, reliable options are PyPDF2 and PyMuPDF (fitz)—neither relies on external tools, so you can keep everything within your Python environment. Here's how to use each:
Option 1: PyPDF2 (Beginner-Friendly & Straightforward)
PyPDF2 is a popular go-to for basic PDF manipulation. It’s easy to learn and works great for simple page extraction tasks.
First, install the library:
pip install PyPDF2
Then use this script to grab pages 2-6 (note: PyPDF2 uses 0-based indexing, so page 2 translates to index 1, page 6 to index 5):
from PyPDF2 import PdfReader, PdfWriter def extract_pdf_pages(input_path, output_path, start_page, end_page): # Open the source PDF reader = PdfReader(input_path) writer = PdfWriter() # Convert 1-based page numbers to 0-based indices start_idx = start_page - 1 end_idx = end_page - 1 # Add each desired page to the output writer for page_num in range(start_idx, end_idx + 1): writer.add_page(reader.pages[page_num]) # Save the final extracted PDF with open(output_path, "wb") as out_file: writer.write(out_file) # Example: Pull pages 2-6 from input.pdf and save as output.pdf extract_pdf_pages("input.pdf", "output.pdf", 2, 6)
Option 2: PyMuPDF (Faster & More Robust)
PyMuPDF (aka fitz) is the better choice if you’re working with large PDFs or need better performance. It handles complex PDFs more smoothly and is significantly faster than PyPDF2.
Install it first:
pip install pymupdf
Here’s the script for extracting pages 2-6 (it also uses 0-based indexing):
import fitz # Import PyMuPDF def extract_pdf_pages(input_path, output_path, start_page, end_page): # Open the source document source_doc = fitz.open(input_path) # Create an empty document for the output output_doc = fitz.open() # Insert the target pages (adjust to 0-based indices) output_doc.insert_pdf(source_doc, from_page=start_page-1, to_page=end_page-1) # Save and clean up output_doc.save(output_path) output_doc.close() source_doc.close() # Example: Extract pages 2-6 from input.pdf to output.pdf extract_pdf_pages("input.pdf", "output.pdf", 2, 6)
Quick Tips:
- Remember both libraries use 0-based indexing—always subtract 1 from your 1-based page numbers (e.g., page 1 = index 0).
- PyMuPDF is the way to go for performance-critical tasks or complex PDFs, while PyPDF2 is simpler for basic use cases.
- No external tools required with either method—everything runs natively in Python.
内容的提问来源于stack exchange,提问作者Nathan Cheval

