You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何利用iTextSharp将PDF文本块Y坐标转换为HTML式从上到下坐标?

解决PDF文本块Y坐标从下到上转从上到下的问题

要把PDF里从下到上的Y坐标转换成HTML/JavaScript/WinForm常用的从上到下的坐标系,核心是利用页面有效高度做反向计算,结合PDF的CropBox(优先)或MediaBox来获取准确的页面范围,下面是具体实现步骤:

核心原理

PDF默认以页面左下角为原点(0,0),Y轴向上递增;而我们需要的坐标系以左上角为原点,Y轴向下递增。转换公式很简单:
转换后的Y值 = 页面有效高度 - PDF原Y值

这里的「页面有效高度」优先取PDF的CropBox(这是实际显示/打印的裁剪区域,是用户能看到的页面范围),如果没有定义CropBox,则用MediaBox(页面的原始尺寸)。

修改你的自定义提取策略

我们需要给LocationTextExtractionStrategyClass添加一个属性来接收页面有效高度,然后在生成文本矩形时自动转换Y坐标:

public class LocationTextExtractionStrategyClass : LocationTextExtractionStrategy {
    // 存储页面有效高度(用于坐标转换)
    public float PageEffectiveHeight { get; set; }
    //Hold each coordinate
    public List<RectAndText> myPoints = new List<RectAndText>();

    //Automatically called for each chunk of text in the PDF
    public override void RenderText(TextRenderInfo renderInfo) {
        base.RenderText(renderInfo);
        var startPosition = 0;
        if (startPosition < 0) {
            return;
        }

        var chars = renderInfo.GetCharacterRenderInfos().ToList();
        var charsText = renderInfo.GetText();
        var firstChar = chars.First();
        var lastChar = chars.Last();

        // 获取PDF原始坐标系下的文本块矩形
        var bottomLeft = firstChar.GetDescentLine().GetStartPoint();
        var topRight = lastChar.GetAscentLine().GetEndPoint();
        var originalRect = new iTextSharp.text.Rectangle(
            bottomLeft[Vector.I1], bottomLeft[Vector.I2],
            topRight[Vector.I1], topRight[Vector.I2]
        );

        // 关键:转换Y坐标到从上到下的坐标系
        float convertedTopY = PageEffectiveHeight - originalRect.Bottom;
        float convertedBottomY = PageEffectiveHeight - originalRect.Top;
        // 生成转换后的矩形
        var convertedRect = new iTextSharp.text.Rectangle(
            originalRect.Left, convertedTopY,
            originalRect.Right, convertedBottomY
        );

        BaseColor curColor = new BaseColor(0f, 0f, 0f);
        if (renderInfo.GetFillColor() != null)
            curColor = renderInfo.GetFillColor();

        // 添加转换后的矩形和文本到集合
        myPoints.Add(new RectAndText(convertedRect, charsText, curColor));
    }
}

使用时获取页面有效高度

在调用提取策略前,先从PDF页面中获取CropBox或MediaBox,计算出页面有效高度:

// 初始化PDF阅读器
using (PdfReader reader = new PdfReader("your_document.pdf")) {
    int targetPageNumber = 1; // 你要处理的页码
    PdfDictionary pageDict = reader.GetPageN(targetPageNumber);

    // 获取页面的CropBox,不存在则用MediaBox
    PdfArray pageBox = pageDict.GetAsArray(PdfName.CROPBOX) ?? pageDict.GetAsArray(PdfName.MEDIABOX);
    if (pageBox == null) {
        // 极端情况:如果都没有定义,默认用A4尺寸(210x297mm转成点:595x842)
        pageBox = new PdfArray(new float[] {0, 0, 595, 842});
    }

    // 计算页面有效高度:PDF坐标中Y轴的最大值减去最小值
    float pageEffectiveHeight = pageBox.GetAsNumber(3).FloatValue - pageBox.GetAsNumber(1).FloatValue;

    // 处理页面旋转(如果有):旋转90/270度时,宽高互换
    int rotation = reader.GetPageRotation(targetPageNumber);
    if (rotation == 90 || rotation == 270) {
        // 旋转后高度等于原宽度
        pageEffectiveHeight = pageBox.GetAsNumber(2).FloatValue - pageBox.GetAsNumber(0).FloatValue;
    }

    // 实例化自定义策略并设置页面高度
    var extractionStrategy = new LocationTextExtractionStrategyClass();
    extractionStrategy.PageEffectiveHeight = pageEffectiveHeight;

    // 提取文本和位置
    PdfTextExtractor.GetTextFromPage(reader, targetPageNumber, extractionStrategy);

    // 现在extractionStrategy.myPoints中的矩形Y坐标就是从上到下的正确值了
    foreach (var item in extractionStrategy.myPoints) {
        // 示例:输出转换后的坐标
        Console.WriteLine($"文本:{item.Text},位置:左{item.Rect.Left},上{item.Rect.Top},右{item.Rect.Right},下{item.Rect.Bottom}");
    }
}

注意事项

  • 页面旋转处理:如果PDF页面有旋转(比如横向文档),一定要先获取旋转角度,旋转90或270度时页面的宽高会互换,需要调整有效高度的计算。
  • 坐标精度:PDF的坐标单位是点(1点=1/72英寸),如果要转成像素,再用你提到的公式像素 = 点 * DPI / 72即可,比如DPI=300时就是像素 = 点 * 300/72。

内容的提问来源于stack exchange,提问作者ell

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 08:33:14