请求提供使用PDFBox提取单词坐标的实现示例
Solution for Getting Word Coordinates with PDFBox
I totally get where you're stuck—extracting individual character positions is a breeze with PDFBox, but rolling those into full words with accurate bounding boxes requires a bit of logic to group adjacent characters properly. Here's a step-by-step implementation that works reliably:
Core Approach
- First, extract all characters along with their individual bounding boxes.
- Group characters into words by checking the horizontal spacing between consecutive characters (if the gap is smaller than a threshold, they belong to the same word).
- For each grouped word, calculate the overall bounding box by taking the minimum x1, minimum y1, maximum x2, and maximum y2 of all characters in the word.
Full Java Implementation
import org.apache.pdfbox.pdmodel.PDDocument; import org.apache.pdfbox.pdmodel.PDPage; import org.apache.pdfbox.text.PDFTextStripper; import org.apache.pdfbox.text.TextPosition; import java.io.File; import java.io.IOException; import java.util.ArrayList; import java.util.List; public class WordCoordinateExtractor extends PDFTextStripper { // Store characters with their positions private List<CharacterPosition> characterPositions = new ArrayList<>(); public WordCoordinateExtractor() throws IOException { super(); } @Override protected void writeString(String text, List<TextPosition> textPositions) throws IOException { for (TextPosition pos : textPositions) { String character = pos.getUnicode(); // Skip whitespace characters (adjust if needed for your use case) if (!character.trim().isEmpty()) { float x1 = pos.getXDirAdj(); float y1 = pos.getYDirAdj(); float x2 = x1 + pos.getWidthDirAdj(); float y2 = y1 - pos.getHeightDirAdj(); // PDF coordinate origin is bottom-left, y increases upward characterPositions.add(new CharacterPosition(character, x1, y1, x2, y2)); } } } // Helper class to hold character and its position private static class CharacterPosition { String character; float x1, y1, x2, y2; public CharacterPosition(String character, float x1, float y1, float x2, float y2) { this.character = character; this.x1 = x1; this.y1 = y1; this.x2 = x2; this.y2 = y2; } } // Group characters into words and calculate word coordinates public List<WordPosition> getWordPositions() { List<WordPosition> wordPositions = new ArrayList<>(); if (characterPositions.isEmpty()) { return wordPositions; } // Start with the first character StringBuilder currentWord = new StringBuilder(characterPositions.get(0).character); float minX = characterPositions.get(0).x1; float minY = characterPositions.get(0).y2; // Since y2 is lower (closer to bottom) float maxX = characterPositions.get(0).x2; float maxY = characterPositions.get(0).y1; // y1 is higher (closer to top) // Threshold for spacing between characters (adjust based on your PDF's font size) float spacingThreshold = 2.0f; for (int i = 1; i < characterPositions.size(); i++) { CharacterPosition currentChar = characterPositions.get(i); CharacterPosition prevChar = characterPositions.get(i - 1); // Check if current character is part of the same word: same line (y difference small) and close enough horizontally boolean sameLine = Math.abs(currentChar.y1 - prevChar.y1) < 5.0f; // Adjust line threshold as needed boolean closeEnough = (currentChar.x1 - prevChar.x2) < spacingThreshold; if (sameLine && closeEnough) { // Add to current word and update bounding box currentWord.append(currentChar.character); minX = Math.min(minX, currentChar.x1); minY = Math.min(minY, currentChar.y2); maxX = Math.max(maxX, currentChar.x2); maxY = Math.max(maxY, currentChar.y1); } else { // Finalize current word wordPositions.add(new WordPosition(currentWord.toString(), minX, minY, maxX, maxY)); // Start new word currentWord = new StringBuilder(currentChar.character); minX = currentChar.x1; minY = currentChar.y2; maxX = currentChar.x2; maxY = currentChar.y1; } } // Add the last word wordPositions.add(new WordPosition(currentWord.toString(), minX, minY, maxX, maxY)); return wordPositions; } // Helper class to hold word and its bounding box public static class WordPosition { String word; float x1, y1, x2, y2; // x1,y1: bottom-left; x2,y2: top-right public WordPosition(String word, float x1, float y1, float x2, float y2) { this.word = word; this.x1 = x1; this.y1 = y1; this.x2 = x2; this.y2 = y2; } @Override public String toString() { return "Word: '" + word + "' | Bounding Box: (" + x1 + ", " + y1 + ") to (" + x2 + ", " + y2 + ")"; } } // Main method to test public static void main(String[] args) { try (PDDocument document = PDDocument.load(new File("your-document.pdf"))) { WordCoordinateExtractor extractor = new WordCoordinateExtractor(); for (PDPage page : document.getPages()) { extractor.characterPositions.clear(); // Reset for each page extractor.extractText(page); List<WordPosition> wordPositions = extractor.getWordPositions(); System.out.println("Page " + (document.getPages().indexOf(page) + 1) + " Word Positions:"); for (WordPosition wp : wordPositions) { System.out.println(wp); } } } catch (IOException e) { e.printStackTrace(); } } }
Key Notes
- Spacing Threshold: The
spacingThresholdvalue (set to 2.0f) might need adjustment based on your PDF's font size and layout. Larger fonts might need a higher threshold. - Line Detection: The
sameLinecheck uses a y-difference threshold of 5.0f—tweak this if your PDF has multi-line text with varying line heights. - Coordinate System: Remember that PDF uses a bottom-left origin, so
y1in theWordPositionclass is the bottom of the word, andy2is the top.
How to Use
- Replace
"your-document.pdf"with your actual PDF file path. - Run the main method—you'll get each word printed along with its bounding box coordinates.
内容的提问来源于stack exchange,提问作者GoodJuJu
相关产品推荐
相关产品推荐

