You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

请求提供使用PDFBox提取单词坐标的实现示例

Solution for Getting Word Coordinates with PDFBox

I totally get where you're stuck—extracting individual character positions is a breeze with PDFBox, but rolling those into full words with accurate bounding boxes requires a bit of logic to group adjacent characters properly. Here's a step-by-step implementation that works reliably:

Core Approach

  1. First, extract all characters along with their individual bounding boxes.
  2. Group characters into words by checking the horizontal spacing between consecutive characters (if the gap is smaller than a threshold, they belong to the same word).
  3. For each grouped word, calculate the overall bounding box by taking the minimum x1, minimum y1, maximum x2, and maximum y2 of all characters in the word.

Full Java Implementation

import org.apache.pdfbox.pdmodel.PDDocument;
import org.apache.pdfbox.pdmodel.PDPage;
import org.apache.pdfbox.text.PDFTextStripper;
import org.apache.pdfbox.text.TextPosition;

import java.io.File;
import java.io.IOException;
import java.util.ArrayList;
import java.util.List;

public class WordCoordinateExtractor extends PDFTextStripper {

    // Store characters with their positions
    private List<CharacterPosition> characterPositions = new ArrayList<>();

    public WordCoordinateExtractor() throws IOException {
        super();
    }

    @Override
    protected void writeString(String text, List<TextPosition> textPositions) throws IOException {
        for (TextPosition pos : textPositions) {
            String character = pos.getUnicode();
            // Skip whitespace characters (adjust if needed for your use case)
            if (!character.trim().isEmpty()) {
                float x1 = pos.getXDirAdj();
                float y1 = pos.getYDirAdj();
                float x2 = x1 + pos.getWidthDirAdj();
                float y2 = y1 - pos.getHeightDirAdj(); // PDF coordinate origin is bottom-left, y increases upward
                characterPositions.add(new CharacterPosition(character, x1, y1, x2, y2));
            }
        }
    }

    // Helper class to hold character and its position
    private static class CharacterPosition {
        String character;
        float x1, y1, x2, y2;

        public CharacterPosition(String character, float x1, float y1, float x2, float y2) {
            this.character = character;
            this.x1 = x1;
            this.y1 = y1;
            this.x2 = x2;
            this.y2 = y2;
        }
    }

    // Group characters into words and calculate word coordinates
    public List<WordPosition> getWordPositions() {
        List<WordPosition> wordPositions = new ArrayList<>();
        if (characterPositions.isEmpty()) {
            return wordPositions;
        }

        // Start with the first character
        StringBuilder currentWord = new StringBuilder(characterPositions.get(0).character);
        float minX = characterPositions.get(0).x1;
        float minY = characterPositions.get(0).y2; // Since y2 is lower (closer to bottom)
        float maxX = characterPositions.get(0).x2;
        float maxY = characterPositions.get(0).y1; // y1 is higher (closer to top)

        // Threshold for spacing between characters (adjust based on your PDF's font size)
        float spacingThreshold = 2.0f;

        for (int i = 1; i < characterPositions.size(); i++) {
            CharacterPosition currentChar = characterPositions.get(i);
            CharacterPosition prevChar = characterPositions.get(i - 1);

            // Check if current character is part of the same word: same line (y difference small) and close enough horizontally
            boolean sameLine = Math.abs(currentChar.y1 - prevChar.y1) < 5.0f; // Adjust line threshold as needed
            boolean closeEnough = (currentChar.x1 - prevChar.x2) < spacingThreshold;

            if (sameLine && closeEnough) {
                // Add to current word and update bounding box
                currentWord.append(currentChar.character);
                minX = Math.min(minX, currentChar.x1);
                minY = Math.min(minY, currentChar.y2);
                maxX = Math.max(maxX, currentChar.x2);
                maxY = Math.max(maxY, currentChar.y1);
            } else {
                // Finalize current word
                wordPositions.add(new WordPosition(currentWord.toString(), minX, minY, maxX, maxY));
                // Start new word
                currentWord = new StringBuilder(currentChar.character);
                minX = currentChar.x1;
                minY = currentChar.y2;
                maxX = currentChar.x2;
                maxY = currentChar.y1;
            }
        }

        // Add the last word
        wordPositions.add(new WordPosition(currentWord.toString(), minX, minY, maxX, maxY));
        return wordPositions;
    }

    // Helper class to hold word and its bounding box
    public static class WordPosition {
        String word;
        float x1, y1, x2, y2; // x1,y1: bottom-left; x2,y2: top-right

        public WordPosition(String word, float x1, float y1, float x2, float y2) {
            this.word = word;
            this.x1 = x1;
            this.y1 = y1;
            this.x2 = x2;
            this.y2 = y2;
        }

        @Override
        public String toString() {
            return "Word: '" + word + "' | Bounding Box: (" + x1 + ", " + y1 + ") to (" + x2 + ", " + y2 + ")";
        }
    }

    // Main method to test
    public static void main(String[] args) {
        try (PDDocument document = PDDocument.load(new File("your-document.pdf"))) {
            WordCoordinateExtractor extractor = new WordCoordinateExtractor();
            for (PDPage page : document.getPages()) {
                extractor.characterPositions.clear(); // Reset for each page
                extractor.extractText(page);
                List<WordPosition> wordPositions = extractor.getWordPositions();

                System.out.println("Page " + (document.getPages().indexOf(page) + 1) + " Word Positions:");
                for (WordPosition wp : wordPositions) {
                    System.out.println(wp);
                }
            }
        } catch (IOException e) {
            e.printStackTrace();
        }
    }
}

Key Notes

  • Spacing Threshold: The spacingThreshold value (set to 2.0f) might need adjustment based on your PDF's font size and layout. Larger fonts might need a higher threshold.
  • Line Detection: The sameLine check uses a y-difference threshold of 5.0f—tweak this if your PDF has multi-line text with varying line heights.
  • Coordinate System: Remember that PDF uses a bottom-left origin, so y1 in the WordPosition class is the bottom of the word, and y2 is the top.

How to Use

  1. Replace "your-document.pdf" with your actual PDF file path.
  2. Run the main method—you'll get each word printed along with its bounding box coordinates.

内容的提问来源于stack exchange,提问作者GoodJuJu

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 09:40:58