You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Selenium WebDriver提取PDF中关联指定文本的目标内容

用Selenium结合PDFTextStripper提取PDF中特定文本对应的金额

嘿,这个需求我之前处理过,其实核心是把Selenium的浏览器操控能力和PDFBox的PDFTextStripper文本读取能力结合起来,咱们一步步来实现:

步骤1:准备依赖

首先得确保你的项目引入了PDFBox的依赖(毕竟PDFTextStripper是它的核心类),如果用Maven的话,在pom.xml里加这段:

<dependency>
    <groupId>org.apache.pdfbox</groupId>
    <artifactId>pdfbox</artifactId>
    <version>2.0.32</version> <!-- 建议用最新稳定版 -->
</dependency>

步骤2:用Selenium处理PDF的访问(可选,针对需要登录/浏览器上下文的场景)

如果你的PDF需要登录后才能查看,或者藏在某个网页链接里,先用Selenium完成前置操作(比如登录、导航到目标页面),然后获取PDF的访问地址:

WebDriver driver = new ChromeDriver();
// 先完成登录或导航到包含PDF的页面
driver.get("https://example.com/your-page-with-pdf");

// 定位PDF链接并获取它的URL
String pdfUrl = driver.findElement(By.linkText("查看账单PDF")).getAttribute("href");

// 如果PDF直接在当前标签页打开,直接拿当前URL就行
// String pdfUrl = driver.getCurrentUrl();

步骤3:读取PDF的全部文本内容

接下来用PDFTextStripper读取PDF的完整文本,这里不需要下载PDF到本地,直接通过URL获取输入流即可:

// 创建URL对象并打开输入流
URL url = new URL(pdfUrl);
InputStream inputStream = url.openStream();

// 加载PDF文档并读取文本
try (PDDocument document = PDDocument.load(inputStream)) {
    PDFTextStripper stripper = new PDFTextStripper();
    String pdfText = stripper.getText(document);
    System.out.println("PDF完整文本:\n" + pdfText);

    // 步骤4:提取目标金额
    // 用正则匹配"Total Monthly Service Charge $xxx.xx"格式的内容
    String regex = "Total Monthly Service Charge \\$(\\d+\\.\\d{2})";
    Pattern pattern = Pattern.compile(regex);
    Matcher matcher = pattern.matcher(pdfText);

    if (matcher.find()) {
        String amount = matcher.group(1);
        System.out.println("提取到的服务费金额:$" + amount);
    } else {
        System.out.println("未找到目标文本对应的金额");
    }
} catch (IOException e) {
    e.printStackTrace();
} finally {
    driver.quit();
}

处理特殊情况:文本换行或格式混乱

如果目标文本和金额不在同一行(比如“Total Monthly Service Charge”单独一行,金额在下一行),调整正则让它能匹配任意空白字符(包括换行):

// 用\\s+匹配空格、换行、制表符等任意空白
String regex = "Total Monthly Service Charge\\s+\\$(\\d+\\.\\d{2})";
// 加上DOTALL模式让.能匹配换行符(可选,根据实际情况调整)
Pattern pattern = Pattern.compile(regex, Pattern.DOTALL);

处理需要权限的PDF

如果PDF需要登录权限,直接用url.openStream()会返回403,这时候把Selenium里的登录Cookie传递给URLConnection就行:

// 从WebDriver获取所有登录Cookie
Set<Cookie> cookies = driver.manage().getCookies();

// 创建URL连接并设置Cookie
URLConnection connection = url.openConnection();
for (Cookie cookie : cookies) {
    connection.addRequestProperty("Cookie", cookie.getName() + "=" + cookie.getValue());
}

// 带着Cookie打开输入流
InputStream inputStream = connection.getInputStream();

内容的提问来源于stack exchange,提问作者Sourabh Roy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 08:38:10