如何在PdfPig中基于Y坐标容差合并文本行?
问题描述
使用PdfPig从订单PDF中读取文本时,原代码会因x.BoundingBox.Bottom的微小差异将单行PDF文本拆分为多行。如何添加容差,使Y坐标存在微小差异的文本归为同一行?原核心代码如下:
static void GetWordsInReadingOrder(UglyToad.PdfPig.Content.Page page, StringBuilder builder) { var words = page.GetWords().ToList(); var wordsList = page.GetWords().GroupBy(x => x.BoundingBox.Bottom); foreach (var word in wordsList) { bool esimene = true; foreach (var item in word) { if (!esimene) builder.Append(' '); esimene = false; builder.Append($"{item.Text}"); } builder.AppendLine(); } }
关键问题出在这行分组逻辑:
var wordsList = page.GetWords().GroupBy(x => x.BoundingBox.Bottom);
我自己编写了一段考虑行起始距离、分配前后排序的代码,想请教这段代码能否优化:
var sorted = (from r in page.GetWords() orderby r.BoundingBox.Bottom, r.BoundingBox.Left select r).ToList(); if (sorted.Count == 0) return; List<WordsWithLines> lines = new List<WordsWithLines>(); const double rowHeight = 10; double lnLastTop = 0; int lnLastRow = 1; double csvReaVpos = sorted.First().BoundingBox.Bottom; foreach (var r in sorted) { if (r.BoundingBox.Bottom >= (lnLastTop + 0.5 * rowHeight) || (r.BoundingBox.Bottom - csvReaVpos) > rowHeight) { ++lnLastRow; csvReaVpos = r.BoundingBox.Bottom; } lnLastTop = r.BoundingBox.Bottom; lines.Add(new WordsWithLines() { RowNo = lnLastRow, WordElement = r }); } var wordsList = from r in lines orderby r.RowNo descending, r.WordElement.BoundingBox.Left select r; int rowNo = wordsList.First().RowNo; bool first = true; foreach (var word in wordsList) { if (rowNo != word.RowNo) { builder.AppendLine(); first = true; rowNo = word.RowNo; } if (!first) builder.Append(' '); builder.Append($"{Core.CharTran(word.WordElement.Text, " ", "_")}"); first = false; }
优化方案与代码改进
核心问题分析
原代码直接用BoundingBox.Bottom分组,PDF中文字的Y坐标可能因排版引擎、字体渲染存在微小浮点误差,导致同一行被拆分成多个分组。你的自定义代码已经尝试用rowHeight作为容差归同行,但逻辑可以更简洁高效。
优化方向
- 简化分组逻辑:用取整或区间分组替代逐值分组,直接处理Y坐标的微小差异
- 减少中间集合:避免创建
WordsWithLines这类中间对象,直接通过排序+分组完成 - 统一排序逻辑:PDF坐标原点通常在左下角,按
Bottom降序才能得到从上到下的阅读顺序(你的代码已处理,但可更清晰)
优化后的极简版本代码
static void GetWordsInReadingOrder(UglyToad.PdfPig.Content.Page page, StringBuilder builder) { var words = page.GetWords().ToList(); if (!words.Any()) return; // 定义容差:根据PDF字体大小调整,比如取字体高度的10%作为容差 const double tolerance = 1.0; // 按Y坐标(Bottom)取整到容差倍数,再按X坐标(Left)排序 var groupedWords = words .GroupBy(w => Math.Round(w.BoundingBox.Bottom / tolerance) * tolerance) .OrderByDescending(g => g.Key) // 从页面顶部到底部排序 .Select(g => g.OrderBy(w => w.BoundingBox.Left)); // 每行内从左到右排序 foreach (var line in groupedWords) { bool isFirstWord = true; foreach (var word in line) { if (!isFirstWord) builder.Append(' '); builder.Append(Core.CharTran(word.Text, " ", "_")); isFirstWord = false; } builder.AppendLine(); } }
针对你自定义代码的优化版本
保留你的行高判断逻辑,简化冗余步骤:
static void GetWordsInReadingOrder(UglyToad.PdfPig.Content.Page page, StringBuilder builder) { var words = page.GetWords().ToList(); if (!words.Any()) return; const double rowHeight = 10; var sortedWords = words.OrderBy(w => w.BoundingBox.Bottom).ThenBy(w => w.BoundingBox.Left).ToList(); int currentRow = 1; double currentRowBaseY = sortedWords[0].BoundingBox.Bottom; var lineGroups = new Dictionary<int, List<UglyToad.PdfPig.Content.Word>>(); lineGroups[currentRow] = new List<UglyToad.PdfPig.Content.Word> { sortedWords[0] }; for (int i = 1; i < sortedWords.Count; i++) { var word = sortedWords[i]; // 判断是否属于当前行:Y差值小于行高的一半 if (Math.Abs(word.BoundingBox.Bottom - currentRowBaseY) > rowHeight * 0.5) { currentRow++; currentRowBaseY = word.BoundingBox.Bottom; lineGroups[currentRow] = new List<UglyToad.PdfPig.Content.Word>(); } lineGroups[currentRow].Add(word); } // 按行号倒序(从页面顶部开始)输出 foreach (var line in lineGroups.OrderByDescending(g => g.Key).Select(g => g.Value)) { bool isFirstWord = true; foreach (var word in line) { if (!isFirstWord) builder.Append(' '); builder.Append(Core.CharTran(word.Text, " ", "_")); isFirstWord = false; } builder.AppendLine(); } }
关键注意事项
- 容差调整:
tolerance或rowHeight需要根据你处理的PDF实际字体大小调整,避免过松或过严 - PDF坐标系统:PdfPig使用左下角为原点的坐标,所以需要按
Bottom降序才能得到从上到下的阅读顺序 - 特殊排版处理:如果PDF存在多列、竖排等复杂排版,需要额外处理X坐标的分组逻辑
内容的提问来源于stack exchange,提问作者Andrus
相关产品推荐
相关产品推荐

