使用wkhtmltopdf生成Sphinx文档PDF时自定义有序TOC的问题
解决Sphinx文档转PDF时wkhtmltopdf目录顺序错乱的问题
我之前在维护Sphinx文档转PDF的流程时,也碰到过一模一样的TOC顺序问题——wkhtmltopdf只会按HTML文件的加载顺序生成目录,完全不管Sphinx的toctree层级逻辑。下面是我试过的两个有效方案,你可以根据自己的场景选择:
方案1:两次生成法(先获取页码再生成自定义TOC)
这个思路是先跑一次生成临时PDF,拿到每个标题对应的实际页码,再结合你已有的toctree结构生成符合预期顺序的自定义TOC页面,最后再生成最终PDF。
步骤1:生成临时PDF并导出目录数据
用python-pdfkit调用wkhtmltopdf时,加上dump-outline选项导出XML格式的原始目录数据:
import pdfkit # 第一次生成临时PDF,同时导出目录XML temp_options = { 'dump-outline': 'outline.xml', # 其他你需要的配置:页面大小、边距、编码等 'page-size': 'A4', 'encoding': 'UTF-8' } # 按原HTML顺序生成临时PDF pdfkit.from_file( ['root/index.html', 'root/child/index.html'], 'temp.pdf', options=temp_options )
步骤2:解析XML获取标题-页码映射
编写脚本解析导出的outline.xml,把每个标题对应的页码存成字典:
import xml.etree.ElementTree as ET def parse_outline(xml_path): tree = ET.parse(xml_path) root = tree.getroot() toc_map = {} # 遍历所有目录项,提取标题和页码 for item in root.iter('item'): title = item.find('title').text page = int(item.find('page').text) toc_map[title] = page return toc_map # 获取标题对应的实际页码 title_page_map = parse_outline('outline.xml')
步骤3:生成自定义TOC页面
用你已有的toctree期望顺序,生成带正确页码的HTML目录页:
# 假设你已通过Sphinx获取的期望TOC顺序 expected_toc = [ 'Heading_1', 'Heading_2', 'Subheading_1', 'Heading_3' ] # 生成自定义TOC的HTML内容 toc_html = """ <!DOCTYPE html> <html> <head> <meta charset="UTF-8"> <title>Table of Contents</title> <style> /* 匹配Sphinx默认样式的TOC样式 */ .toc-container { margin: 2em auto; max-width: 80%; } .toc { list-style: none; padding-left: 0; } .toc > li { margin: 1em 0; font-size: 1.1em; } .toc ul { padding-left: 1.5em; margin-top: 0.5em; } .toc a { color: #404040; text-decoration: none; } .toc a:hover { text-decoration: underline; } </style> </head> <body> <div class="toc-container"> <h1>Table of Contents</h1> <ul class="toc"> """ # 按期望顺序生成TOC条目 for title in expected_toc: if title in title_page_map: # 根据标题层级调整缩进(比如子标题嵌套) if title == 'Subheading_1': toc_html += ' <ul>\n <li><a href="#{}">{}</a> (Page {})</li>\n </ul>\n'.format( title.lower().replace(' ', '_'), title, title_page_map[title] ) else: toc_html += ' <li><a href="#{}">{}</a> (Page {})</li>\n'.format( title.lower().replace(' ', '_'), title, title_page_map[title] ) toc_html += """ </ul> </div> </body> </html> """ # 保存自定义TOC页面 with open('custom_toc.html', 'w', encoding='utf-8') as f: f.write(toc_html)
步骤4:生成最终PDF
把自定义TOC页面放在最前面,关闭wkhtmltopdf的默认TOC选项,生成最终PDF:
final_options = { 'no-outline': None, # 禁用默认TOC 'page-size': 'A4', 'encoding': 'UTF-8' } # 注意顺序:先加载自定义TOC,再加载文档HTML pdfkit.from_file( ['custom_toc.html', 'root/index.html', 'root/child/index.html'], 'final.pdf', options=final_options )
方案2:合并HTML页面,让内容按toctree顺序流式输出
如果不想跑两次生成流程,可以在Sphinx生成HTML后,把子页面的内容直接插入到主页面的对应位置,让整个文档变成一个连续的HTML流,这样wkhtmltopdf自然会按内容顺序生成正确的TOC。
示例合并脚本
def merge_sphinx_html(main_html_path, child_html_path, insert_after_heading): # 读取主页面内容 with open(main_html_path, 'r', encoding='utf-8') as f: main_content = f.read() # 读取子页面内容,提取正文部分(去掉head和body标签) with open(child_html_path, 'r', encoding='utf-8') as f: child_content = f.read() child_body_start = child_content.find('<body>') + 6 child_body_end = child_content.rfind('</body>') child_body = child_content[child_body_start:child_body_end] # 在主页面的指定标题后插入子页面内容 # 注意:这里的标签要和Sphinx实际生成的标题标签一致(比如h1/h2) insert_marker = f'<h1>{insert_after_heading}</h1>' if insert_marker in main_content: merged_content = main_content.replace(insert_marker, insert_marker + '\n' + child_body) with open('merged_index.html', 'w', encoding='utf-8') as f: f.write(merged_content) else: print(f"未找到标题 {insert_after_heading},无法合并") # 把child页面内容插入到主页面的Heading_2之后 merge_sphinx_html('root/index.html', 'root/child/index.html', 'Heading_2')
然后直接转换合并后的HTML即可:
pdfkit.from_file('merged_index.html', 'final.pdf', options={'page-size': 'A4'})
注意事项
- 方案1中要确保标题文本完全匹配,Sphinx生成的标题可能带有细微格式差异(比如特殊字符),需要调整解析逻辑适配。
- 方案2中合并HTML时,要注意子页面的样式、脚本是否和主页面兼容,避免出现样式错乱。
内容的提问来源于stack exchange,提问作者Begoña Álvarez de la Cruz
相关产品推荐
相关产品推荐

