You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用PDFBox的PDFTextStripper提取PDF中的特定数据?

Solutions to Extract Specific Strings (like price or age) with PDFBox

Hey there! Since you’ve already got full text extraction working with PDFTextStripper, targeting specific strings like price or age just boils down to adding targeted text matching logic—either on the fully extracted text, or by hooking into PDFBox’s low-level text processing pipeline. Here are practical, actionable approaches:

1. Simple String/Regex Matching on Extracted Text

If your PDF has relatively structured text, this is the fastest way to get results. Once you’ve pulled the full text (or line-by-line content) using PDFTextStripper, use Java’s built-in string tools or regular expressions to hunt for your target terms and their associated values.

Example: Regex to Extract Price Values

Suppose you want to grab numerical values paired with "price" (accounting for case variations and different separators like colons or spaces):

import java.util.regex.Matcher;
import java.util.regex.Pattern;

// After extracting full text with PDFTextStripper
String pdfText = stripper.getText(document);

// Regex pattern: matches "price" (any case), optional spaces/colons, then a number
Pattern pricePattern = Pattern.compile("price\\s*[:\\s]*([\\d\\.]+)", Pattern.CASE_INSENSITIVE);
Matcher matcher = pricePattern.matcher(pdfText);

// Iterate through all matches and extract the value
while (matcher.find()) {
    String priceValue = matcher.group(1);
    System.out.println("Found price: " + priceValue);
}

You can adapt this pattern for age (e.g., age\\s*[:\\s]*([\\d]+) ) or any other term—tweak the regex to fit your PDF’s specific formatting (like currency symbols or hyphens).

2. Custom PDFTextStripper for Real-Time Word Monitoring

If you need more control (like tracking the on-page position of keywords, or handling non-continuous text chunks), extend PDFTextStripper and override the processTextPosition method. This lets you inspect every text fragment as PDFBox parses it.

Example: Custom Stripper to Track Keywords and Context

import org.apache.pdfbox.text.PDFTextStripper;
import org.apache.pdfbox.text.TextPosition;
import java.io.IOException;
import java.util.ArrayList;
import java.util.List;

public class KeywordStripper extends PDFTextStripper {
    private final List<String> targetKeywords = List.of("price", "age");
    private final List<String> foundMatches = new ArrayList<>();
    private final StringBuilder currentWord = new StringBuilder();
    private boolean captureNextWord = false;

    public KeywordStripper() throws IOException {
        super();
    }

    @Override
    protected void processTextPosition(TextPosition text) {
        String character = text.getUnicode();
        // Build words by detecting whitespace (adjust logic for your PDF's spacing behavior)
        if (Character.isWhitespace(character.charAt(0))) {
            String word = currentWord.toString().toLowerCase();
            if (targetKeywords.contains(word)) {
                foundMatches.add("Found keyword: " + word);
                captureNextWord = true; // Optional: Flag to capture the following value
            } else if (captureNextWord) {
                foundMatches.add("Corresponding value: " + word);
                captureNextWord = false;
            }
            currentWord.setLength(0);
        } else {
            currentWord.append(character);
        }
        super.processTextPosition(text);
    }

    public List<String> getFoundMatches() {
        return foundMatches;
    }
}

Use it like this:

KeywordStripper stripper = new KeywordStripper();
stripper.getText(document);
List<String> matches = stripper.getFoundMatches();

This approach is perfect if you need to link keywords to their exact on-page coordinates (via text.getX()/text.getY()) or handle tricky formatting where words are split across multiple PDF text chunks.

3. Handle Complex Layouts (e.g., Tables)

If price or age live in tables or non-linear layouts, you’ll need to group text by position:

  • Group text chunks by their Y-coordinate (round to a small decimal to account for minor position differences) to identify rows
  • For each assembled row, check if it contains your keyword, then extract the corresponding value from the same row

You can extend the processTextPosition method to track rows and build row-level text strings, then apply regex matching per row for precise extraction.

Key Tips for Success

  • Case Insensitivity: Always use Pattern.CASE_INSENSITIVE or normalize text to lowercase/uppercase to avoid missing matches like Price or PRICE.
  • Flexible Regex: Account for different separators (e.g., price: 100, price $100, price - 100) by using regex groups and optional characters.
  • Test with Your PDFs: PDF formatting varies wildly—test your logic against your specific documents to adjust spacing, regex, or word-building rules.

内容的提问来源于stack exchange,提问作者Ismail Brahmi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 07:51:27