Python如何获取Firefox内置PDF阅读器打开PDF对应的完整HTML代码
你当前获取不到完整内容的核心原因有两个:
- Firefox内置PDF阅读器基于PDF.js实现,正文内容默认绘制在
<canvas>元素上,不会以明文文本DOM节点的形式存在于初始页面源码中 - 即使开启了文本选择层,PDF.js默认采用懒加载策略,仅会渲染当前视口附近页面的文本节点,初始加载时仅生成侧边栏缩略图、页码相关的DOM结构
最稳定的实现方案是直接调用PDF.js暴露在全局的PDFViewerApplication内置接口获取内容,不需要解析页面源码,文本、坐标、字体、样式等信息都可以直接提取,信息丰富度远高于PyPDF的解析结果。
修改后的可运行代码如下:
from selenium.webdriver.firefox.service import Service from selenium.webdriver.firefox.options import Options from selenium.webdriver.support.ui import WebDriverWait from selenium import webdriver import os service = Service(os.path.abspath('Files/geckodriver')) options = Options() # 无头模式可正常运行,无需开启图形界面 options.headless = True driver = webdriver.Firefox(service = service, options = options) driver.get(f'file://{os.path.abspath("Files/sample.pdf")}') # 等待PDF.js和文档加载完成 wait = WebDriverWait(driver, 15) wait.until(lambda d: d.execute_script(""" return typeof PDFViewerApplication !== 'undefined' && PDFViewerApplication.pdfDocument !== null """)) # 调用JS接口批量获取所有页面的文本及附加信息 pdf_data = driver.execute_script(""" return new Promise((resolve) => { const totalPages = PDFViewerApplication.pdfDocument.numPages; const result = []; let loadedCount = 0; for (let pageNum = 1; pageNum <= totalPages; pageNum++) { PDFViewerApplication.pdfDocument.getPage(pageNum).then(page => { page.getTextContent().then(content => { // 可以根据需求提取content内的坐标、字体、字号等信息 const pageText = content.items.map(item => item.str).join(' '); result.push({ page_num: pageNum, text: pageText, raw_items: content.items // 保留所有结构化信息,不需要可以删掉 }); loadedCount++; if (loadedCount === totalPages) resolve(result); }); }); } }); """) # 输出结果 for page in pdf_data: print(f"===== 第{page['page_num']}页 =====") print(page['text']) driver.quit()
如果确实需要获取渲染后的完整HTML源码,需要额外配置PDF.js关闭懒加载、强制渲染所有页面的文本层,再模拟滚动触发全量页面加载后再调用page_source,但该方案运行效率低、稳定性差,不如直接调用内置API。
内容的提问来源于stack exchange,提问作者Random User
相关产品推荐
相关产品推荐

