You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用wkhtmltopdf生成Sphinx文档PDF时自定义有序TOC的问题

解决Sphinx文档转PDF时wkhtmltopdf目录顺序错乱的问题

我之前在维护Sphinx文档转PDF的流程时,也碰到过一模一样的TOC顺序问题——wkhtmltopdf只会按HTML文件的加载顺序生成目录,完全不管Sphinx的toctree层级逻辑。下面是我试过的两个有效方案,你可以根据自己的场景选择:

方案1:两次生成法(先获取页码再生成自定义TOC)

这个思路是先跑一次生成临时PDF,拿到每个标题对应的实际页码,再结合你已有的toctree结构生成符合预期顺序的自定义TOC页面,最后再生成最终PDF。

步骤1:生成临时PDF并导出目录数据

用python-pdfkit调用wkhtmltopdf时,加上dump-outline选项导出XML格式的原始目录数据:

import pdfkit

# 第一次生成临时PDF,同时导出目录XML
temp_options = {
    'dump-outline': 'outline.xml',
    # 其他你需要的配置:页面大小、边距、编码等
    'page-size': 'A4',
    'encoding': 'UTF-8'
}

# 按原HTML顺序生成临时PDF
pdfkit.from_file(
    ['root/index.html', 'root/child/index.html'],
    'temp.pdf',
    options=temp_options
)

步骤2:解析XML获取标题-页码映射

编写脚本解析导出的outline.xml,把每个标题对应的页码存成字典:

import xml.etree.ElementTree as ET

def parse_outline(xml_path):
    tree = ET.parse(xml_path)
    root = tree.getroot()
    toc_map = {}
    # 遍历所有目录项,提取标题和页码
    for item in root.iter('item'):
        title = item.find('title').text
        page = int(item.find('page').text)
        toc_map[title] = page
    return toc_map

# 获取标题对应的实际页码
title_page_map = parse_outline('outline.xml')

步骤3:生成自定义TOC页面

用你已有的toctree期望顺序,生成带正确页码的HTML目录页:

# 假设你已通过Sphinx获取的期望TOC顺序
expected_toc = [
    'Heading_1',
    'Heading_2',
    'Subheading_1',
    'Heading_3'
]

# 生成自定义TOC的HTML内容
toc_html = """
<!DOCTYPE html>
<html>
<head>
    <meta charset="UTF-8">
    <title>Table of Contents</title>
    <style>
        /* 匹配Sphinx默认样式的TOC样式 */
        .toc-container { margin: 2em auto; max-width: 80%; }
        .toc { list-style: none; padding-left: 0; }
        .toc > li { margin: 1em 0; font-size: 1.1em; }
        .toc ul { padding-left: 1.5em; margin-top: 0.5em; }
        .toc a { color: #404040; text-decoration: none; }
        .toc a:hover { text-decoration: underline; }
    </style>
</head>
<body>
    <div class="toc-container">
        <h1>Table of Contents</h1>
        <ul class="toc">
"""

# 按期望顺序生成TOC条目
for title in expected_toc:
    if title in title_page_map:
        # 根据标题层级调整缩进(比如子标题嵌套)
        if title == 'Subheading_1':
            toc_html += '            <ul>\n                <li><a href="#{}">{}</a> (Page {})</li>\n            </ul>\n'.format(
                title.lower().replace(' ', '_'), title, title_page_map[title]
            )
        else:
            toc_html += '            <li><a href="#{}">{}</a> (Page {})</li>\n'.format(
                title.lower().replace(' ', '_'), title, title_page_map[title]
            )

toc_html += """
        </ul>
    </div>
</body>
</html>
"""

# 保存自定义TOC页面
with open('custom_toc.html', 'w', encoding='utf-8') as f:
    f.write(toc_html)

步骤4:生成最终PDF

把自定义TOC页面放在最前面,关闭wkhtmltopdf的默认TOC选项,生成最终PDF:

final_options = {
    'no-outline': None,  # 禁用默认TOC
    'page-size': 'A4',
    'encoding': 'UTF-8'
}

# 注意顺序:先加载自定义TOC,再加载文档HTML
pdfkit.from_file(
    ['custom_toc.html', 'root/index.html', 'root/child/index.html'],
    'final.pdf',
    options=final_options
)

方案2:合并HTML页面,让内容按toctree顺序流式输出

如果不想跑两次生成流程,可以在Sphinx生成HTML后,把子页面的内容直接插入到主页面的对应位置,让整个文档变成一个连续的HTML流,这样wkhtmltopdf自然会按内容顺序生成正确的TOC。

示例合并脚本

def merge_sphinx_html(main_html_path, child_html_path, insert_after_heading):
    # 读取主页面内容
    with open(main_html_path, 'r', encoding='utf-8') as f:
        main_content = f.read()
    # 读取子页面内容,提取正文部分(去掉head和body标签)
    with open(child_html_path, 'r', encoding='utf-8') as f:
        child_content = f.read()
    child_body_start = child_content.find('<body>') + 6
    child_body_end = child_content.rfind('</body>')
    child_body = child_content[child_body_start:child_body_end]
    
    # 在主页面的指定标题后插入子页面内容
    # 注意:这里的标签要和Sphinx实际生成的标题标签一致(比如h1/h2)
    insert_marker = f'<h1>{insert_after_heading}</h1>'
    if insert_marker in main_content:
        merged_content = main_content.replace(insert_marker, insert_marker + '\n' + child_body)
        with open('merged_index.html', 'w', encoding='utf-8') as f:
            f.write(merged_content)
    else:
        print(f"未找到标题 {insert_after_heading},无法合并")

# 把child页面内容插入到主页面的Heading_2之后
merge_sphinx_html('root/index.html', 'root/child/index.html', 'Heading_2')

然后直接转换合并后的HTML即可:

pdfkit.from_file('merged_index.html', 'final.pdf', options={'page-size': 'A4'})

注意事项

  • 方案1中要确保标题文本完全匹配,Sphinx生成的标题可能带有细微格式差异(比如特殊字符),需要调整解析逻辑适配。
  • 方案2中合并HTML时,要注意子页面的样式、脚本是否和主页面兼容,避免出现样式错乱。

内容的提问来源于stack exchange,提问作者Begoña Álvarez de la Cruz

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.12 04:39:00