You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用PyPDF2完整复制PDF时仅复制内容未复制大纲的问题求助

解决PyPDF2复制PDF时丢失大纲(书签)的问题

嘿,我之前也踩过这个坑!PyPDF2 默认只复制页面内容,原文档的大纲(书签)不会自动同步到新文件里,得手动把书签结构搬过去才行。下面是修改后的完整代码,能帮你完美复制PDF的内容和大纲:

import sys
from PyPDF2 import PdfReader, PdfWriter

def add_outline_items(outline_items, parent=None, writer=None):
    """递归添加大纲项到新PDF"""
    for item in outline_items:
        # 添加当前书签,关联对应页码
        bookmark = writer.add_outline_item(
            title=item.title,
            page_number=item.page_number,
            parent=parent
        )
        # 如果当前书签有子项,递归添加嵌套书签
        if item.children:
            add_outline_items(item.children, parent=bookmark, writer=writer)

def copy_pdf_with_outline(input_path, output_path):
    # 读取原PDF文件
    reader = PdfReader(input_path)
    writer = PdfWriter()

    # 先复制所有页面到新文档
    for page in reader.pages:
        writer.add_page(page)

    # 检查原文档是否有大纲,有就复制过去
    if reader.outline:
        add_outline_items(reader.outline, writer=writer)

    # 保存最终的PDF文件
    with open(output_path, "wb") as out_file:
        writer.write(out_file)

if __name__ == "__main__":
    # 校验命令行参数是否正确
    if len(sys.argv) != 3:
        print("使用方式: python test.py <input pdf> <output dest>")
        sys.exit(1)
    input_pdf = sys.argv[1]
    output_dest = sys.argv[2]
    copy_pdf_with_outline(input_pdf, output_dest)
    print(f"PDF已成功复制到 {output_dest},包含原文档大纲!")

关键细节说明:

  • add_outline_items 是核心函数:因为PDF的大纲可能有多层嵌套(比如一级标题下的二级子标题),所以用递归的方式遍历所有大纲项,确保每个层级的书签都能正确添加。
  • 先复制页面再处理大纲:这样能保证书签关联的页码和新文档的页码完全匹配,不会出现跳转错误。
  • 增加了参数校验:避免用户因为输入参数错误导致程序报错。

使用方法:

直接用你原来的命令运行就行:

python test.py 你的输入文件.pdf 目标输出文件.pdf

这样生成的新PDF就会和原文档完全一致,内容和大纲都不会丢失啦!

内容的提问来源于stack exchange,提问作者sp00kyb00g13

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 03:26:07