使用Selenium打印含MathJax的网页为PDF时公式空白问题求助
解决Selenium打印PDF时MathJax空白及整合wkhtmltopdf的方案
一、修复Selenium打印时MathJax空白的问题
你的问题根源在于两个点:一是MathJax渲染检测脚本适配性差,二是time.sleep无法精准等待渲染完成,再加上Firefox打印默认设置没适配MathJax渲染规则。
1. 替换MathJax渲染等待脚本
不同版本MathJax的API差异大,你用的MathJax.Hub是旧版(2.x)写法,现在很多网站用3.x,所以要做兼容,并且要等渲染完全结束再打印:
from selenium.webdriver.support.ui import WebDriverWait def wait_for_mathjax(): # 兼容MathJax 2.x和3.x的渲染等待脚本 wait_script = """ return new Promise(resolve => { if (typeof MathJax !== 'undefined') { if (MathJax.Hub) { // MathJax 2.x MathJax.Hub.Queue(["Typeset", MathJax.Hub], () => { console.log('MathJax 2.x渲染完成'); resolve(true); }); } else if (MathJax.typesetPromise) { // MathJax 3.x MathJax.typesetPromise().then(() => { console.log('MathJax 3.x渲染完成'); resolve(true); }).catch(err => { console.log('MathJax渲染出错:', err); resolve(true); }); } else { console.log('MathJax已渲染或版本不兼容'); resolve(true); } } else { console.log('未检测到MathJax'); resolve(true); } }); """ # 用WebDriverWait替代time.sleep,精准等待渲染完成 WebDriverWait(driver, 30).until(lambda d: d.execute_script(wait_script)) def print_page(): # 先等页面基础元素加载完成 WebDriverWait(driver, 20).until(EC.presence_of_element_located((By.TAG_NAME, "body"))) # 等待MathJax渲染完成 wait_for_mathjax() # 配置Firefox打印参数(关键:启用背景,避免MathJax元素被截断) print_settings = { "recentDestinations": [{ "id": "Save as PDF", "origin": "local", "account": "" }], "selectedDestinationId": "Save as PDF", "version": 2, "isHeaderFooterEnabled": False, "isBackgroundEnabled": True } driver.execute_script("window.print();") time.sleep(5)
2. 优化Firefox初始化配置
启动driver时添加打印相关的profile设置,确保PDF生成时包含MathJax内容:
options = Options() profile = webdriver.FirefoxProfile() profile.set_preference("print.print_to_file", True) profile.set_preference("print.save_as_pdf.images.enabled", True) profile.set_preference("print.save_as_pdf.links.enabled", True) options.profile = profile driver = webdriver.Firefox(options=options)
二、整合wkhtmltopdf到现有流程
既然wkhtmltopdf能正常生成带MathJax的PDF,你可以利用Selenium已登录的状态,导出cookie给wkhtmltopdf,不用重复登录就能打印目标页面。
1. 从Selenium导出cookie
登录完成后,把浏览器cookie转换成wkhtmltopdf能识别的格式:
def get_wkhtml_cookies(): cookies = driver.get_cookies() cookie_args = [] for cookie in cookies: cookie_str = f"{cookie['name']}={cookie['value']}" cookie_args.append(f"--cookie {cookie_str}") return " ".join(cookie_args)
2. 调用wkhtmltopdf生成PDF
用Python的subprocess模块执行wkhtmltopdf命令,传入cookie和目标URL:
import subprocess def print_with_wkhtml(url, output_path): cookie_args = get_wkhtml_cookies() # 构建命令,添加参数确保MathJax渲染完成 cmd = f"wkhtmltopdf {cookie_args} --enable-javascript --javascript-delay 10000 --no-stop-slow-scripts {url} {output_path}" result = subprocess.run(cmd, shell=True, capture_output=True, text=True) if result.returncode == 0: print(f"PDF生成成功:{output_path}") else: print(f"生成失败:{result.stderr}")
3. 使用示例
登录完成后,遍历需要打印的页面:
# 假设已完成登录流程 target_urls = ["https://example.com/page1", "https://example.com/page2"] for idx, url in enumerate(target_urls): driver.get(url) WebDriverWait(driver, 20).until(EC.presence_of_element_located((By.TAG_NAME, "body"))) output_path = f"page_{idx+1}.pdf" print_with_wkhtml(url, output_path)
注意事项
- 先在PopOS上安装wkhtmltopdf:执行
sudo apt install wkhtmltopdf即可。 - 如果页面MathJax加载慢,可把
--javascript-delay的数值调大(单位是毫秒),比如改成15000。 - 导出cookie要在登录后执行,确保cookie有效。
内容的提问来源于stack exchange,提问作者P E
相关产品推荐
相关产品推荐

