如何利用iTextSharp将PDF文本块Y坐标转换为HTML式从上到下坐标?
解决PDF文本块Y坐标从下到上转从上到下的问题
要把PDF里从下到上的Y坐标转换成HTML/JavaScript/WinForm常用的从上到下的坐标系,核心是利用页面有效高度做反向计算,结合PDF的CropBox(优先)或MediaBox来获取准确的页面范围,下面是具体实现步骤:
核心原理
PDF默认以页面左下角为原点(0,0),Y轴向上递增;而我们需要的坐标系以左上角为原点,Y轴向下递增。转换公式很简单:转换后的Y值 = 页面有效高度 - PDF原Y值
这里的「页面有效高度」优先取PDF的CropBox(这是实际显示/打印的裁剪区域,是用户能看到的页面范围),如果没有定义CropBox,则用MediaBox(页面的原始尺寸)。
修改你的自定义提取策略
我们需要给LocationTextExtractionStrategyClass添加一个属性来接收页面有效高度,然后在生成文本矩形时自动转换Y坐标:
public class LocationTextExtractionStrategyClass : LocationTextExtractionStrategy { // 存储页面有效高度(用于坐标转换) public float PageEffectiveHeight { get; set; } //Hold each coordinate public List<RectAndText> myPoints = new List<RectAndText>(); //Automatically called for each chunk of text in the PDF public override void RenderText(TextRenderInfo renderInfo) { base.RenderText(renderInfo); var startPosition = 0; if (startPosition < 0) { return; } var chars = renderInfo.GetCharacterRenderInfos().ToList(); var charsText = renderInfo.GetText(); var firstChar = chars.First(); var lastChar = chars.Last(); // 获取PDF原始坐标系下的文本块矩形 var bottomLeft = firstChar.GetDescentLine().GetStartPoint(); var topRight = lastChar.GetAscentLine().GetEndPoint(); var originalRect = new iTextSharp.text.Rectangle( bottomLeft[Vector.I1], bottomLeft[Vector.I2], topRight[Vector.I1], topRight[Vector.I2] ); // 关键:转换Y坐标到从上到下的坐标系 float convertedTopY = PageEffectiveHeight - originalRect.Bottom; float convertedBottomY = PageEffectiveHeight - originalRect.Top; // 生成转换后的矩形 var convertedRect = new iTextSharp.text.Rectangle( originalRect.Left, convertedTopY, originalRect.Right, convertedBottomY ); BaseColor curColor = new BaseColor(0f, 0f, 0f); if (renderInfo.GetFillColor() != null) curColor = renderInfo.GetFillColor(); // 添加转换后的矩形和文本到集合 myPoints.Add(new RectAndText(convertedRect, charsText, curColor)); } }
使用时获取页面有效高度
在调用提取策略前,先从PDF页面中获取CropBox或MediaBox,计算出页面有效高度:
// 初始化PDF阅读器 using (PdfReader reader = new PdfReader("your_document.pdf")) { int targetPageNumber = 1; // 你要处理的页码 PdfDictionary pageDict = reader.GetPageN(targetPageNumber); // 获取页面的CropBox,不存在则用MediaBox PdfArray pageBox = pageDict.GetAsArray(PdfName.CROPBOX) ?? pageDict.GetAsArray(PdfName.MEDIABOX); if (pageBox == null) { // 极端情况:如果都没有定义,默认用A4尺寸(210x297mm转成点:595x842) pageBox = new PdfArray(new float[] {0, 0, 595, 842}); } // 计算页面有效高度:PDF坐标中Y轴的最大值减去最小值 float pageEffectiveHeight = pageBox.GetAsNumber(3).FloatValue - pageBox.GetAsNumber(1).FloatValue; // 处理页面旋转(如果有):旋转90/270度时,宽高互换 int rotation = reader.GetPageRotation(targetPageNumber); if (rotation == 90 || rotation == 270) { // 旋转后高度等于原宽度 pageEffectiveHeight = pageBox.GetAsNumber(2).FloatValue - pageBox.GetAsNumber(0).FloatValue; } // 实例化自定义策略并设置页面高度 var extractionStrategy = new LocationTextExtractionStrategyClass(); extractionStrategy.PageEffectiveHeight = pageEffectiveHeight; // 提取文本和位置 PdfTextExtractor.GetTextFromPage(reader, targetPageNumber, extractionStrategy); // 现在extractionStrategy.myPoints中的矩形Y坐标就是从上到下的正确值了 foreach (var item in extractionStrategy.myPoints) { // 示例:输出转换后的坐标 Console.WriteLine($"文本:{item.Text},位置:左{item.Rect.Left},上{item.Rect.Top},右{item.Rect.Right},下{item.Rect.Bottom}"); } }
注意事项
- 页面旋转处理:如果PDF页面有旋转(比如横向文档),一定要先获取旋转角度,旋转90或270度时页面的宽高会互换,需要调整有效高度的计算。
- 坐标精度:PDF的坐标单位是点(1点=1/72英寸),如果要转成像素,再用你提到的公式
像素 = 点 * DPI / 72即可,比如DPI=300时就是像素 = 点 * 300/72。
内容的提问来源于stack exchange,提问作者ell
相关产品推荐
相关产品推荐

