You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python的fitz库拆分PDF并保留对应页范围的目录?

拆分PDF并保留对应页范围的导航目录

我有一本800多页的PDF书籍,带有仅用于阅读器导航的目录(并非PDF页面内容),需要拆分该PDF,使每个拆分后的文件包含对应页范围的目录,功能类似ilovepdf的拆分工具。已知PyMuPDF(fitz)模块提供Document.get_toc()和Document.set_toc()用于获取、设置目录,但以下代码仅能完成PDF拆分,无法保留对应目录:

import fitz

def split_pdf_with_toc(input_pdf_path, output_folder, page_ranges):
    # Open the PDF file
    pdf_document = fitz.open(input_pdf_path)

    for i, (start_page, end_page) in enumerate(page_ranges, start=1):
        # Create a new PDF for each specified page range
        chapter_pdf = fitz.open()
        chapter_pdf.insert_pdf(pdf_document, from_page=start_page - 1, to_page=end_page - 1)

        # Save the new PDF with the corresponding TOC
        output_pdf_path = f"{output_folder}/Chapter_{i}.pdf"
        chapter_pdf.save(output_pdf_path)
        chapter_pdf.close()

    # Close the original PDF
    pdf_document.close()

# Example usage
input_pdf_path = "demo/robin.pdf"
output_folder = "output"
# Specify your desired page ranges as a list of tuples
page_ranges = [(45, 81)]  # Add more tuples as needed

split_pdf_with_toc(input_pdf_path, output_folder, page_ranges)

解决方案:保留目录的PDF拆分代码

修改原代码,加入目录筛选与页码调整逻辑,确保拆分后的PDF包含对应页范围的导航目录:

import fitz
import os

def split_pdf_with_toc(input_pdf_path, output_folder, page_ranges):
    # 创建输出目录(如果不存在)
    os.makedirs(output_folder, exist_ok=True)
    
    # 打开原PDF并获取完整目录
    pdf_document = fitz.open(input_pdf_path)
    original_toc = pdf_document.get_toc()

    for i, (start_page, end_page) in enumerate(page_ranges, start=1):
        # 创建新PDF并插入指定页范围
        chapter_pdf = fitz.open()
        chapter_pdf.insert_pdf(pdf_document, from_page=start_page - 1, to_page=end_page - 1)

        # 筛选并调整当前拆分范围对应的目录项
        filtered_toc = []
        for entry in original_toc:
            level, title, page_num = entry
            # 仅保留页码落在当前拆分范围内的目录项
            if start_page <= page_num <= end_page:
                # 调整页码:新PDF页码从1开始,计算偏移后的页码
                new_page_num = page_num - start_page + 1
                filtered_toc.append([level, title, new_page_num])
        
        # 将调整后的目录设置到新PDF中
        chapter_pdf.set_toc(filtered_toc)

        # 保存拆分后的PDF
        output_pdf_path = f"{output_folder}/Chapter_{i}.pdf"
        chapter_pdf.save(output_pdf_path)
        chapter_pdf.close()

    pdf_document.close()

# 示例调用
input_pdf_path = "demo/robin.pdf"
output_folder = "output"
# 页范围为(起始页, 结束页),支持多个范围
page_ranges = [(45, 81), (82, 150)]

split_pdf_with_toc(input_pdf_path, output_folder, page_ranges)

关键逻辑说明

  • 获取原目录:通过get_toc()获取原PDF的完整导航目录,目录项格式为[层级, 标题, 页码]
  • 筛选目录项:遍历原目录,仅保留页码落在当前拆分页范围内的条目
  • 调整页码:拆分后的PDF页码从1重新计数,因此需要将原目录页码减去拆分起始页,再加1得到新页码
  • 设置新目录:使用set_toc()将调整后的目录应用到新生成的PDF文件中

内容的提问来源于stack exchange,提问作者Sabbir Ahmed

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.05 06:32:53