You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python如何获取Firefox内置PDF阅读器打开PDF对应的完整HTML代码

你当前获取不到完整内容的核心原因有两个:

  • Firefox内置PDF阅读器基于PDF.js实现,正文内容默认绘制在<canvas>元素上,不会以明文文本DOM节点的形式存在于初始页面源码中
  • 即使开启了文本选择层,PDF.js默认采用懒加载策略,仅会渲染当前视口附近页面的文本节点,初始加载时仅生成侧边栏缩略图、页码相关的DOM结构

最稳定的实现方案是直接调用PDF.js暴露在全局的PDFViewerApplication内置接口获取内容,不需要解析页面源码,文本、坐标、字体、样式等信息都可以直接提取,信息丰富度远高于PyPDF的解析结果。

修改后的可运行代码如下:

from selenium.webdriver.firefox.service import Service
from selenium.webdriver.firefox.options import Options
from selenium.webdriver.support.ui import WebDriverWait
from selenium import webdriver
import os

service = Service(os.path.abspath('Files/geckodriver'))

options = Options()
# 无头模式可正常运行,无需开启图形界面
options.headless = True

driver = webdriver.Firefox(service = service, options = options)
driver.get(f'file://{os.path.abspath("Files/sample.pdf")}')

# 等待PDF.js和文档加载完成
wait = WebDriverWait(driver, 15)
wait.until(lambda d: d.execute_script("""
    return typeof PDFViewerApplication !== 'undefined' 
        && PDFViewerApplication.pdfDocument !== null
"""))

# 调用JS接口批量获取所有页面的文本及附加信息
pdf_data = driver.execute_script("""
    return new Promise((resolve) => {
        const totalPages = PDFViewerApplication.pdfDocument.numPages;
        const result = [];
        let loadedCount = 0;
        
        for (let pageNum = 1; pageNum <= totalPages; pageNum++) {
            PDFViewerApplication.pdfDocument.getPage(pageNum).then(page => {
                page.getTextContent().then(content => {
                    // 可以根据需求提取content内的坐标、字体、字号等信息
                    const pageText = content.items.map(item => item.str).join(' ');
                    result.push({
                        page_num: pageNum,
                        text: pageText,
                        raw_items: content.items // 保留所有结构化信息,不需要可以删掉
                    });
                    loadedCount++;
                    if (loadedCount === totalPages) resolve(result);
                });
            });
        }
    });
""")

# 输出结果
for page in pdf_data:
    print(f"===== 第{page['page_num']}页 =====")
    print(page['text'])

driver.quit()

如果确实需要获取渲染后的完整HTML源码,需要额外配置PDF.js关闭懒加载、强制渲染所有页面的文本层,再模拟滚动触发全量页面加载后再调用page_source,但该方案运行效率低、稳定性差,不如直接调用内置API。

内容的提问来源于stack exchange,提问作者Random User

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.24 16:36:01