You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

iTextSharp 5.x能否获取PDF中二维码及文本的坐标?

Answer to Your PDF Coordinate Extraction Question

Hey there! Great question—let me break this down for you based on my experience working with PDF libraries, especially iTextSharp 5.x.

Can iTextSharp 5.x get coordinates for QR codes and specific text?

Absolutely, though the approach differs a bit for text vs. QR codes:

1. Extracting specific text coordinates

To grab the (x,y) position of a target text string, you'll need to use a custom text extraction strategy that tracks each text chunk's location. iTextSharp's built-in LocationTextExtractionStrategy can be extended to capture this data.

Here's a simplified example:

public class TextLocationStrategy : LocationTextExtractionStrategy
{
    public List<TextChunkData> TextLocations { get; } = new List<TextChunkData>();

    public override void RenderText(TextRenderInfo renderInfo)
    {
        base.RenderText(renderInfo);
        var startPoint = renderInfo.GetBaseline().GetStartPoint();
        // Store text and its coordinates (note: PDF origin is bottom-left corner)
        TextLocations.Add(new TextChunkData(renderInfo.GetText(), startPoint[0], startPoint[1]));
    }
}

public class TextChunkData
{
    public string Text { get; }
    public float X { get; }
    public float Y { get; }

    public TextChunkData(string text, float x, float y)
    {
        Text = text;
        X = x;
        Y = y;
    }
}

// Usage in your code:
using (var reader = new PdfReader("your-file.pdf"))
{
    for (int page = 1; page <= reader.NumberOfPages; page++)
    {
        var strategy = new TextLocationStrategy();
        PdfTextExtractor.GetTextFromPage(reader, page, strategy);
        
        // Find your target text
        var targetChunk = strategy.TextLocations.FirstOrDefault(t => t.Text.Contains("your-target-string"));
        if (targetChunk != null)
        {
            Console.WriteLine($"Target text found at X: {targetChunk.X}, Y: {targetChunk.Y}");
        }
    }
}

Keep in mind: PDF uses a coordinate system where (0,0) sits at the bottom-left corner of the page, not the top-left like most screen-based systems.

2. Extracting QR code coordinates

QR codes in PDFs are typically stored as either raster images or vector graphics. For raster images, you can parse the page's resources to locate image XObjects and calculate their position:

using (var reader = new PdfReader("your-file.pdf"))
{
    for (int page = 1; page <= reader.NumberOfPages; page++)
    {
        var pageDict = reader.GetPageN(page);
        var resources = pageDict.GetAsDict(PdfName.RESOURCES);
        var xObjects = resources?.GetAsDict(PdfName.XOBJECT);
        
        if (xObjects != null)
        {
            foreach (var key in xObjects.Keys)
            {
                var xObj = xObjects.GetAsIndirectObject(key);
                var xObjDict = PdfReader.GetPdfObject(xObj) as PdfDictionary;
                
                if (xObjDict?.GetAsName(PdfName.SUBTYPE) == PdfName.IMAGE)
                {
                    // Get the transformation matrix to find the image's bottom-left corner
                    var matrix = PdfReader.GetPdfObject(xObjDict.GetAsArray(PdfName.MATRIX)) as PdfArray;
                    if (matrix != null)
                    {
                        // Matrix values follow [a, b, c, d, e, f] where (e,f) is the origin point
                        float x = matrix.GetAsNumber(4).FloatValue;
                        float y = matrix.GetAsNumber(5).FloatValue;
                        Console.WriteLine($"Image (likely QR code) found at X: {x}, Y: {y}");
                    }
                }
            }
        }
    }
}

If the QR code is vector-based (generated as path graphics), this method won't work—iTextSharp 5.x doesn't natively detect vector QR code patterns, so you'd need a specialized barcode parser for that scenario.

Alternative frameworks if iTextSharp 5.x falls short

If you hit limitations (like vector QR codes or more complex coordinate logic), these libraries are solid alternatives:

  • iText 7: The modern successor to iTextSharp, with a cleaner API and better support for advanced PDF features. It simplifies coordinate extraction and image detection significantly.
  • Apache PDFBox: A free, open-source library (available for C# via IKVM). It has robust tools for extracting text bounding boxes and image positions, with straightforward APIs.
  • Spire.PDF: A commercial library with a free tier, offering built-in barcode/QR code detection and intuitive methods to grab coordinates for text and images.

内容的提问来源于stack exchange,提问作者Ninos

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 08:56:34