在C# .NET中使用PDFBox获取PDF特定区域文字的问题求助
Got it, extracting text from a specific region (like your x=20, y=30, width=100, height=100 area) in PDFBox is totally manageable—here are two reliable approaches you can use, depending on your needs:
Approach 1: Use PDFTextStripperByArea (Simplest Method)
This class is built specifically for extracting text from defined regions, making it the easiest way to get what you need. Just note that PDF uses a bottom-left origin—if your coordinates are measured from the top of the page, you’ll need to convert the Y value first.
Example Code
import org.apache.pdfbox.pdmodel.PDDocument; import org.apache.pdfbox.text.PDFTextStripperByArea; import org.apache.pdfbox.pdmodel.common.PDRectangle; import java.io.File; import java.io.IOException; public class RegionTextExtractor { public static void main(String[] args) { String pdfFilePath = "your-document.pdf"; // Your target region (adjust Y if using top-origin coordinates) float targetX = 20; float targetY = 30; float targetWidth = 100; float targetHeight = 100; try (PDDocument document = PDDocument.load(new File(pdfFilePath))) { PDFTextStripperByArea stripper = new PDFTextStripperByArea(); stripper.setSortByPosition(true); // Optional: keeps text in reading order // Convert top-origin Y to PDF's bottom-origin var page = document.getPage(0); float pageHeight = page.getMediaBox().getHeight(); float adjustedY = pageHeight - targetY - targetHeight; // Define and add your target region PDRectangle targetRegion = new PDRectangle(targetX, adjustedY, targetWidth, targetHeight); stripper.addRegion("myTargetRegion", targetRegion); // Extract text from the page stripper.extractRegions(page); String extractedText = stripper.getTextForRegion("myTargetRegion"); System.out.println("Extracted text from target region:\n" + extractedText); } catch (IOException e) { e.printStackTrace(); } } }
Approach 2: Custom PDFTextStripper (More Control)
If you need fine-grained control over how text segments are filtered, you can extend PDFTextStripper and override the writeString method to check each text position against your region.
Example Code
import org.apache.pdfbox.pdmodel.PDDocument; import org.apache.pdfbox.text.PDFTextStripper; import org.apache.pdfbox.text.TextPosition; import java.io.File; import java.io.IOException; import java.util.List; public class BoundedTextStripper extends PDFTextStripper { private final float targetX; private final float targetY; private final float targetWidth; private final float targetHeight; public BoundedTextStripper(float x, float y, float width, float height) throws IOException { super(); this.targetX = x; this.targetY = y; this.targetWidth = width; this.targetHeight = height; } @Override protected void writeString(String text, List<TextPosition> textPositions) throws IOException { for (TextPosition textPos : textPositions) { // Get the bounding box of the current text segment float textX = textPos.getX(); float textY = textPos.getY(); float textW = textPos.getWidth(); float textH = textPos.getHeight(); // Check if the text segment intersects our target region boolean isInRegion = (textX + textW >= targetX) && (textX <= targetX + targetWidth) && (textY + textH >= targetY) && (textY <= targetY + targetHeight); if (isInRegion) { getOutputWriter().write(textPos.getUnicode()); } } } // Usage public static void main(String[] args) { String pdfPath = "your-document.pdf"; float x = 20; float y = 30; float w = 100; float h = 100; try (PDDocument doc = PDDocument.load(new File(pdfPath))) { // Convert top-origin Y to bottom-origin if needed var page = doc.getPage(0); float pageHeight = page.getMediaBox().getHeight(); float adjustedY = pageHeight - y - h; BoundedTextStripper stripper = new BoundedTextStripper(x, adjustedY, w, h); stripper.setStartPage(0); stripper.setEndPage(0); String extractedText = stripper.getText(doc); System.out.println("Extracted text:\n" + extractedText); } catch (IOException e) { e.printStackTrace(); } } }
Key Notes to Avoid Headaches
- Coordinate System: Always double-check your origin! PDF uses bottom-left, but most UI tools (like PDF viewers) use top-left. Use the
adjustedYcalculation above to convert if needed. - Scanned PDFs: PDFBox only extracts selectable text. If your PDF is a scanned image, you’ll need to run OCR first (try combining PDFBox with Tesseract).
- Text Segmentation: Some PDFs split text into small segments. Using the intersection check (instead of strict "contains") ensures you don’t miss partial text in your region.
- Dependencies: Make sure you’re using the latest stable PDFBox version. For Maven, add this to your
pom.xml:<dependency> <groupId>org.apache.pdfbox</groupId> <artifactId>pdfbox</artifactId> <version>2.0.32</version> <!-- Use the latest stable release --> </dependency>
内容的提问来源于stack exchange,提问作者Rabiul Aleem

