You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在C# .NET中使用PDFBox获取PDF特定区域文字的问题求助

Extracting Text from a Specific Bounding Box with PDFBox

Got it, extracting text from a specific region (like your x=20, y=30, width=100, height=100 area) in PDFBox is totally manageable—here are two reliable approaches you can use, depending on your needs:

Approach 1: Use PDFTextStripperByArea (Simplest Method)

This class is built specifically for extracting text from defined regions, making it the easiest way to get what you need. Just note that PDF uses a bottom-left origin—if your coordinates are measured from the top of the page, you’ll need to convert the Y value first.

Example Code

import org.apache.pdfbox.pdmodel.PDDocument;
import org.apache.pdfbox.text.PDFTextStripperByArea;
import org.apache.pdfbox.pdmodel.common.PDRectangle;
import java.io.File;
import java.io.IOException;

public class RegionTextExtractor {
    public static void main(String[] args) {
        String pdfFilePath = "your-document.pdf";
        
        // Your target region (adjust Y if using top-origin coordinates)
        float targetX = 20;
        float targetY = 30;
        float targetWidth = 100;
        float targetHeight = 100;

        try (PDDocument document = PDDocument.load(new File(pdfFilePath))) {
            PDFTextStripperByArea stripper = new PDFTextStripperByArea();
            stripper.setSortByPosition(true); // Optional: keeps text in reading order

            // Convert top-origin Y to PDF's bottom-origin
            var page = document.getPage(0);
            float pageHeight = page.getMediaBox().getHeight();
            float adjustedY = pageHeight - targetY - targetHeight;

            // Define and add your target region
            PDRectangle targetRegion = new PDRectangle(targetX, adjustedY, targetWidth, targetHeight);
            stripper.addRegion("myTargetRegion", targetRegion);

            // Extract text from the page
            stripper.extractRegions(page);
            String extractedText = stripper.getTextForRegion("myTargetRegion");

            System.out.println("Extracted text from target region:\n" + extractedText);
        } catch (IOException e) {
            e.printStackTrace();
        }
    }
}

Approach 2: Custom PDFTextStripper (More Control)

If you need fine-grained control over how text segments are filtered, you can extend PDFTextStripper and override the writeString method to check each text position against your region.

Example Code

import org.apache.pdfbox.pdmodel.PDDocument;
import org.apache.pdfbox.text.PDFTextStripper;
import org.apache.pdfbox.text.TextPosition;
import java.io.File;
import java.io.IOException;
import java.util.List;

public class BoundedTextStripper extends PDFTextStripper {
    private final float targetX;
    private final float targetY;
    private final float targetWidth;
    private final float targetHeight;

    public BoundedTextStripper(float x, float y, float width, float height) throws IOException {
        super();
        this.targetX = x;
        this.targetY = y;
        this.targetWidth = width;
        this.targetHeight = height;
    }

    @Override
    protected void writeString(String text, List<TextPosition> textPositions) throws IOException {
        for (TextPosition textPos : textPositions) {
            // Get the bounding box of the current text segment
            float textX = textPos.getX();
            float textY = textPos.getY();
            float textW = textPos.getWidth();
            float textH = textPos.getHeight();

            // Check if the text segment intersects our target region
            boolean isInRegion = (textX + textW >= targetX) && 
                                 (textX <= targetX + targetWidth) && 
                                 (textY + textH >= targetY) && 
                                 (textY <= targetY + targetHeight);

            if (isInRegion) {
                getOutputWriter().write(textPos.getUnicode());
            }
        }
    }

    // Usage
    public static void main(String[] args) {
        String pdfPath = "your-document.pdf";
        float x = 20;
        float y = 30;
        float w = 100;
        float h = 100;

        try (PDDocument doc = PDDocument.load(new File(pdfPath))) {
            // Convert top-origin Y to bottom-origin if needed
            var page = doc.getPage(0);
            float pageHeight = page.getMediaBox().getHeight();
            float adjustedY = pageHeight - y - h;

            BoundedTextStripper stripper = new BoundedTextStripper(x, adjustedY, w, h);
            stripper.setStartPage(0);
            stripper.setEndPage(0);

            String extractedText = stripper.getText(doc);
            System.out.println("Extracted text:\n" + extractedText);
        } catch (IOException e) {
            e.printStackTrace();
        }
    }
}

Key Notes to Avoid Headaches

  • Coordinate System: Always double-check your origin! PDF uses bottom-left, but most UI tools (like PDF viewers) use top-left. Use the adjustedY calculation above to convert if needed.
  • Scanned PDFs: PDFBox only extracts selectable text. If your PDF is a scanned image, you’ll need to run OCR first (try combining PDFBox with Tesseract).
  • Text Segmentation: Some PDFs split text into small segments. Using the intersection check (instead of strict "contains") ensures you don’t miss partial text in your region.
  • Dependencies: Make sure you’re using the latest stable PDFBox version. For Maven, add this to your pom.xml:
    <dependency>
        <groupId>org.apache.pdfbox</groupId>
        <artifactId>pdfbox</artifactId>
        <version>2.0.32</version> <!-- Use the latest stable release -->
    </dependency>
    

内容的提问来源于stack exchange,提问作者Rabiul Aleem

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 07:24:49