You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PDFBox 2.0.8提取可见文本时部分单词首字母缺失问题及方案咨询

Fixing Partial First Character Extraction When Filtering Visible PDF Text with Custom PageDrawer

Let me break down what's likely going wrong and how to fix that partial "Tena..."首字母 issue you're seeing:

Why the Partial Character Problem Happens

When you copied the PageDrawer code to check clipping paths, you're probably relying on the bounding box (BBox) of the text position to judge visibility, rather than the actual rendered glyph path. Here's the catch:

  • Some characters (especially those with stylized or extended shapes, or positioned with text matrix offsets) have a BBox that doesn't perfectly match their visible glyph. For example, the first letter of "Tena..." might have a small portion of its BBox falling outside the clipping path, even though the visible part of the glyph is fully inside.
  • Or, you might not be correctly applying all transformation matrices (text matrix + current transformation matrix) when checking the glyph against the clipping path, leading to coordinate mismatches.

Step-by-Step Fixes

1. Prioritize Glyph Paths Over Bounding Boxes

Instead of checking if the text's BBox is inside the clipping path, use the actual rendered glyph path for visibility checks. This gives you precise control over what's truly visible.

Here's how to adjust your custom PageDrawer or PdfTextStripper subclass:

private boolean isGlyphVisible(TextPosition textPos) {
    GraphicsState gs = getGraphicsState();
    // Combine the current transformation matrix (CTM) and text matrix (Tm)
    AffineTransform combinedTransform = new AffineTransform(gs.getCurrentTransformationMatrix());
    combinedTransform.concatenate(textPos.getTextMatrix());

    // Get the actual glyph path (not just the BBox)
    Path glyphPath = textPos.getGlyphPath();
    if (glyphPath == null) {
        // Fallback to BBox if glyph path isn't available
        Rectangle2D bbox = textPos.getBoundingBox();
        Shape transformedBBox = combinedTransform.createTransformedShape(bbox);
        return gs.getClippingPath().intersects(transformedBBox.getBounds2D());
    }

    // Transform the glyph path to user space
    Shape transformedGlyph = combinedTransform.createTransformedShape(glyphPath);
    // Check if any part of the glyph intersects the clipping path
    return gs.getClippingPath().intersects(transformedGlyph.getBounds2D()) 
        || gs.getClippingPath().contains(transformedGlyph.getBounds2D());
}

2. Fix Transformation Matrix Handling

Make sure you're concatenating matrices in the correct order: CTM first, then text matrix. This ensures the glyph coordinates are correctly mapped to the same user space as the clipping path. Skipping this step is a common cause of false "invisible" judgments for edge characters.

3. Adjust Partial Visibility Logic

If a glyph is partially visible (like your truncated first letter), you probably still want to extract the full character instead of discarding it. The code above uses intersects() instead of contains(), so even a small overlap between the glyph and clipping path will mark the character as visible.

4. Avoid Full PageDrawer Copy (Optional)

You don't need to duplicate the entire PageDrawer code. Instead, subclass PdfTextStripper and override writeString() to filter visible characters directly:

public class VisibleTextStripper extends PdfTextStripper {
    public VisibleTextStripper() throws IOException {
        super();
    }

    @Override
    protected void writeString(String text, List<TextPosition> textPositions) throws IOException {
        StringBuilder visibleText = new StringBuilder();
        for (TextPosition textPos : textPositions) {
            if (isGlyphVisible(textPos)) {
                visibleText.append(textPos.getUnicode());
            }
        }
        super.writeString(visibleText.toString(), textPositions);
    }

    // Include the isGlyphVisible() method from above here
}

Additional Tips

  • Check PDFBox Version: Ensure you're using a recent version of PDFBox (2.0+), as older versions have bugs in glyph path retrieval and matrix handling.
  • Performance Optimization: For large PDFs, add a quick BBox check first to filter out obviously invisible characters, then run the precise glyph path check only on candidates. This balances accuracy and speed.
  • Test Edge Cases: For the "Tena..." example, inspect the PDF's text matrix and clipping path using a tool like PDFDebugger (part of PDFBox) to confirm the coordinate mismatch.

内容的提问来源于stack exchange,提问作者D.F. Stones

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 07:27:14