Selenium读取URL中PDF内容时PDDocument load方法报错如何解决
问题排查与解决方案
问题1:The method load(BufferedInputStream) is undefined for the type PDDocument报错
原因
该报错为PDFBox版本API不匹配或类导入错误导致:
- 若使用PDFBox 3.0及以上版本,官方已将原静态
load()方法重命名为open() - 若使用2.x版本仍报错,要么是导入了非Apache PDFBox的同名
PDDocument类,要么是依赖包版本不一致(pdfbox和fontbox版本必须完全对应)
解决步骤
- 首先确认导包正确:
import org.apache.pdfbox.pdmodel.PDDocument; import org.apache.pdfbox.text.PDFTextStripper; import java.io.BufferedInputStream; import java.io.InputStream; import java.net.URL; import java.io.IOException;
- 若使用3.x版本,将加载代码替换为:
PDDocument doc = PDDocument.open(bf);
- 推荐使用稳定2.x版本,Maven依赖示例:
<dependency> <groupId>org.apache.pdfbox</groupId> <artifactId>pdfbox</artifactId> <version>2.0.32</version> </dependency> <dependency> <groupId>org.apache.pdfbox</groupId> <artifactId>fontbox</artifactId> <version>2.0.32</version> </dependency>
问题2:无法从URL加载PDF
原因
你提供的URL是PDF在线预览页面地址,返回的是HTML内容,不是PDF二进制文件流,直接请求无法解析为PDF;另外该接口需要身份校验,直接用URL.openStream()没有携带浏览器的登录Cookie,会被服务器拒绝访问。
解决步骤
- 替换为PDF直链:找到站点返回PDF二进制流的直接链接,而不是预览页地址
- 携带身份信息请求:如果你用Selenium操作浏览器已经登录,可以从Selenium实例中获取所有Cookie,添加到PDF请求的头信息中,示例如下:
// 从Selenium driver获取Cookie,添加到请求头 Map<String, String> cookies = driver.manage().getCookies().stream() .collect(Collectors.toMap(Cookie::getName, Cookie::getValue)); String cookieStr = cookies.entrySet().stream() .map(e -> e.getKey() + "=" + e.getValue()) .collect(Collectors.joining("; ")); String userAgent = driver.executeScript("return navigator.userAgent").toString();
- 优化资源管理:用try-with-resources语法自动关闭流和PDDocument资源,避免内存泄漏,不需要手动调用
close()方法。
修正后的完整工具方法
public static String readPdfContent(String pdfDirectUrl, String cookieStr, String userAgent) throws IOException { URL pdfUrl = new URL(pdfDirectUrl); HttpURLConnection conn = (HttpURLConnection) pdfUrl.openConnection(); conn.setRequestProperty("Cookie", cookieStr); conn.setRequestProperty("User-Agent", userAgent); try (InputStream in = conn.getInputStream(); BufferedInputStream bf = new BufferedInputStream(in); PDDocument doc = PDDocument.load(bf)) { int numberOfPages = doc.getNumberOfPages(); System.out.println("The total number of pages " + numberOfPages); return new PDFTextStripper().getText(doc); } }
内容的提问来源于stack exchange,提问作者ankush singh
相关产品推荐
相关产品推荐

