PDFBox 2.0.8提取可见文本时部分单词首字母缺失问题及方案咨询
Let me break down what's likely going wrong and how to fix that partial "Tena..."首字母 issue you're seeing:
Why the Partial Character Problem Happens
When you copied the PageDrawer code to check clipping paths, you're probably relying on the bounding box (BBox) of the text position to judge visibility, rather than the actual rendered glyph path. Here's the catch:
- Some characters (especially those with stylized or extended shapes, or positioned with text matrix offsets) have a BBox that doesn't perfectly match their visible glyph. For example, the first letter of "Tena..." might have a small portion of its BBox falling outside the clipping path, even though the visible part of the glyph is fully inside.
- Or, you might not be correctly applying all transformation matrices (text matrix + current transformation matrix) when checking the glyph against the clipping path, leading to coordinate mismatches.
Step-by-Step Fixes
1. Prioritize Glyph Paths Over Bounding Boxes
Instead of checking if the text's BBox is inside the clipping path, use the actual rendered glyph path for visibility checks. This gives you precise control over what's truly visible.
Here's how to adjust your custom PageDrawer or PdfTextStripper subclass:
private boolean isGlyphVisible(TextPosition textPos) { GraphicsState gs = getGraphicsState(); // Combine the current transformation matrix (CTM) and text matrix (Tm) AffineTransform combinedTransform = new AffineTransform(gs.getCurrentTransformationMatrix()); combinedTransform.concatenate(textPos.getTextMatrix()); // Get the actual glyph path (not just the BBox) Path glyphPath = textPos.getGlyphPath(); if (glyphPath == null) { // Fallback to BBox if glyph path isn't available Rectangle2D bbox = textPos.getBoundingBox(); Shape transformedBBox = combinedTransform.createTransformedShape(bbox); return gs.getClippingPath().intersects(transformedBBox.getBounds2D()); } // Transform the glyph path to user space Shape transformedGlyph = combinedTransform.createTransformedShape(glyphPath); // Check if any part of the glyph intersects the clipping path return gs.getClippingPath().intersects(transformedGlyph.getBounds2D()) || gs.getClippingPath().contains(transformedGlyph.getBounds2D()); }
2. Fix Transformation Matrix Handling
Make sure you're concatenating matrices in the correct order: CTM first, then text matrix. This ensures the glyph coordinates are correctly mapped to the same user space as the clipping path. Skipping this step is a common cause of false "invisible" judgments for edge characters.
3. Adjust Partial Visibility Logic
If a glyph is partially visible (like your truncated first letter), you probably still want to extract the full character instead of discarding it. The code above uses intersects() instead of contains(), so even a small overlap between the glyph and clipping path will mark the character as visible.
4. Avoid Full PageDrawer Copy (Optional)
You don't need to duplicate the entire PageDrawer code. Instead, subclass PdfTextStripper and override writeString() to filter visible characters directly:
public class VisibleTextStripper extends PdfTextStripper { public VisibleTextStripper() throws IOException { super(); } @Override protected void writeString(String text, List<TextPosition> textPositions) throws IOException { StringBuilder visibleText = new StringBuilder(); for (TextPosition textPos : textPositions) { if (isGlyphVisible(textPos)) { visibleText.append(textPos.getUnicode()); } } super.writeString(visibleText.toString(), textPositions); } // Include the isGlyphVisible() method from above here }
Additional Tips
- Check PDFBox Version: Ensure you're using a recent version of PDFBox (2.0+), as older versions have bugs in glyph path retrieval and matrix handling.
- Performance Optimization: For large PDFs, add a quick BBox check first to filter out obviously invisible characters, then run the precise glyph path check only on candidates. This balances accuracy and speed.
- Test Edge Cases: For the "Tena..." example, inspect the PDF's text matrix and clipping path using a tool like PDFDebugger (part of PDFBox) to confirm the coordinate mismatch.
内容的提问来源于stack exchange,提问作者D.F. Stones

