You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Selenium读取URL中PDF内容时PDDocument load方法报错如何解决

问题排查与解决方案

问题1:The method load(BufferedInputStream) is undefined for the type PDDocument报错

原因

该报错为PDFBox版本API不匹配或类导入错误导致:

  • 若使用PDFBox 3.0及以上版本,官方已将原静态load()方法重命名为open()
  • 若使用2.x版本仍报错,要么是导入了非Apache PDFBox的同名PDDocument类,要么是依赖包版本不一致(pdfbox和fontbox版本必须完全对应)

解决步骤

  • 首先确认导包正确:
import org.apache.pdfbox.pdmodel.PDDocument;
import org.apache.pdfbox.text.PDFTextStripper;
import java.io.BufferedInputStream;
import java.io.InputStream;
import java.net.URL;
import java.io.IOException;
  • 若使用3.x版本,将加载代码替换为:
PDDocument doc = PDDocument.open(bf);
  • 推荐使用稳定2.x版本,Maven依赖示例:
<dependency>
    <groupId>org.apache.pdfbox</groupId>
    <artifactId>pdfbox</artifactId>
    <version>2.0.32</version>
</dependency>
<dependency>
    <groupId>org.apache.pdfbox</groupId>
    <artifactId>fontbox</artifactId>
    <version>2.0.32</version>
</dependency>

问题2:无法从URL加载PDF

原因

你提供的URL是PDF在线预览页面地址,返回的是HTML内容,不是PDF二进制文件流,直接请求无法解析为PDF;另外该接口需要身份校验,直接用URL.openStream()没有携带浏览器的登录Cookie,会被服务器拒绝访问。

解决步骤

  1. 替换为PDF直链:找到站点返回PDF二进制流的直接链接,而不是预览页地址
  2. 携带身份信息请求:如果你用Selenium操作浏览器已经登录,可以从Selenium实例中获取所有Cookie,添加到PDF请求的头信息中,示例如下:
// 从Selenium driver获取Cookie,添加到请求头
Map<String, String> cookies = driver.manage().getCookies().stream()
    .collect(Collectors.toMap(Cookie::getName, Cookie::getValue));
String cookieStr = cookies.entrySet().stream()
    .map(e -> e.getKey() + "=" + e.getValue())
    .collect(Collectors.joining("; "));
String userAgent = driver.executeScript("return navigator.userAgent").toString();
  1. 优化资源管理:用try-with-resources语法自动关闭流和PDDocument资源,避免内存泄漏,不需要手动调用close()方法。

修正后的完整工具方法

public static String readPdfContent(String pdfDirectUrl, String cookieStr, String userAgent) throws IOException {
    URL pdfUrl = new URL(pdfDirectUrl);
    HttpURLConnection conn = (HttpURLConnection) pdfUrl.openConnection();
    conn.setRequestProperty("Cookie", cookieStr);
    conn.setRequestProperty("User-Agent", userAgent);

    try (InputStream in = conn.getInputStream();
         BufferedInputStream bf = new BufferedInputStream(in);
         PDDocument doc = PDDocument.load(bf)) {
        int numberOfPages = doc.getNumberOfPages();
        System.out.println("The total number of pages " + numberOfPages);
        return new PDFTextStripper().getText(doc);
    }
}

内容的提问来源于stack exchange,提问作者ankush singh

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.04 07:12:00