You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在PdfPig中基于Y坐标容差合并文本行?

问题描述

使用PdfPig从订单PDF中读取文本时,原代码会因x.BoundingBox.Bottom的微小差异将单行PDF文本拆分为多行。如何添加容差,使Y坐标存在微小差异的文本归为同一行?原核心代码如下:

static void GetWordsInReadingOrder(UglyToad.PdfPig.Content.Page page, StringBuilder builder)
{ 
  var words = page.GetWords().ToList();
  var wordsList = page.GetWords().GroupBy(x => x.BoundingBox.Bottom);
  foreach (var word in wordsList)
  {
    bool esimene = true;
    foreach (var item in word)
    {
      if (!esimene)
        builder.Append(' ');
      esimene = false;
      builder.Append($"{item.Text}");
    }
    builder.AppendLine();
  }
}

关键问题出在这行分组逻辑:

var wordsList = page.GetWords().GroupBy(x => x.BoundingBox.Bottom);

我自己编写了一段考虑行起始距离、分配前后排序的代码,想请教这段代码能否优化:

var sorted = (from r in page.GetWords()
              orderby r.BoundingBox.Bottom, r.BoundingBox.Left
              select r).ToList();

if (sorted.Count == 0)
  return;
List<WordsWithLines> lines = new List<WordsWithLines>();
const double rowHeight = 10;
double lnLastTop = 0;
int lnLastRow = 1;
double csvReaVpos = sorted.First().BoundingBox.Bottom;
foreach (var r in sorted)
{
  if (r.BoundingBox.Bottom >= (lnLastTop + 0.5 * rowHeight) || (r.BoundingBox.Bottom  - csvReaVpos) > rowHeight)
  {
    ++lnLastRow;
    csvReaVpos = r.BoundingBox.Bottom;
  }
  lnLastTop = r.BoundingBox.Bottom;
  lines.Add(new WordsWithLines()
  {
    RowNo = lnLastRow,
    WordElement = r
  });
}
var wordsList = from r in lines
                orderby r.RowNo descending, r.WordElement.BoundingBox.Left
                select r;

int rowNo = wordsList.First().RowNo;
bool first = true;
foreach (var word in wordsList)
{
  if (rowNo != word.RowNo)
  {
    builder.AppendLine();
    first = true;
    rowNo = word.RowNo;
  }
  if (!first)  
      builder.Append(' ');
    builder.Append($"{Core.CharTran(word.WordElement.Text, " ", "_")}");
  first = false;
}

优化方案与代码改进

核心问题分析

原代码直接用BoundingBox.Bottom分组,PDF中文字的Y坐标可能因排版引擎、字体渲染存在微小浮点误差,导致同一行被拆分成多个分组。你的自定义代码已经尝试用rowHeight作为容差归同行,但逻辑可以更简洁高效。

优化方向

  1. 简化分组逻辑:用取整或区间分组替代逐值分组,直接处理Y坐标的微小差异
  2. 减少中间集合:避免创建WordsWithLines这类中间对象,直接通过排序+分组完成
  3. 统一排序逻辑:PDF坐标原点通常在左下角,按Bottom降序才能得到从上到下的阅读顺序(你的代码已处理,但可更清晰)

优化后的极简版本代码

static void GetWordsInReadingOrder(UglyToad.PdfPig.Content.Page page, StringBuilder builder)
{
    var words = page.GetWords().ToList();
    if (!words.Any()) return;

    // 定义容差:根据PDF字体大小调整,比如取字体高度的10%作为容差
    const double tolerance = 1.0;
    // 按Y坐标(Bottom)取整到容差倍数,再按X坐标(Left)排序
    var groupedWords = words
        .GroupBy(w => Math.Round(w.BoundingBox.Bottom / tolerance) * tolerance)
        .OrderByDescending(g => g.Key) // 从页面顶部到底部排序
        .Select(g => g.OrderBy(w => w.BoundingBox.Left)); // 每行内从左到右排序

    foreach (var line in groupedWords)
    {
        bool isFirstWord = true;
        foreach (var word in line)
        {
            if (!isFirstWord)
                builder.Append(' ');
            builder.Append(Core.CharTran(word.Text, " ", "_"));
            isFirstWord = false;
        }
        builder.AppendLine();
    }
}

针对你自定义代码的优化版本

保留你的行高判断逻辑,简化冗余步骤:

static void GetWordsInReadingOrder(UglyToad.PdfPig.Content.Page page, StringBuilder builder)
{
    var words = page.GetWords().ToList();
    if (!words.Any()) return;

    const double rowHeight = 10;
    var sortedWords = words.OrderBy(w => w.BoundingBox.Bottom).ThenBy(w => w.BoundingBox.Left).ToList();

    int currentRow = 1;
    double currentRowBaseY = sortedWords[0].BoundingBox.Bottom;
    var lineGroups = new Dictionary<int, List<UglyToad.PdfPig.Content.Word>>();
    lineGroups[currentRow] = new List<UglyToad.PdfPig.Content.Word> { sortedWords[0] };

    for (int i = 1; i < sortedWords.Count; i++)
    {
        var word = sortedWords[i];
        // 判断是否属于当前行:Y差值小于行高的一半
        if (Math.Abs(word.BoundingBox.Bottom - currentRowBaseY) > rowHeight * 0.5)
        {
            currentRow++;
            currentRowBaseY = word.BoundingBox.Bottom;
            lineGroups[currentRow] = new List<UglyToad.PdfPig.Content.Word>();
        }
        lineGroups[currentRow].Add(word);
    }

    // 按行号倒序(从页面顶部开始)输出
    foreach (var line in lineGroups.OrderByDescending(g => g.Key).Select(g => g.Value))
    {
        bool isFirstWord = true;
        foreach (var word in line)
        {
            if (!isFirstWord)
                builder.Append(' ');
            builder.Append(Core.CharTran(word.Text, " ", "_"));
            isFirstWord = false;
        }
        builder.AppendLine();
    }
}

关键注意事项

  • 容差调整:tolerance或rowHeight需要根据你处理的PDF实际字体大小调整,避免过松或过严
  • PDF坐标系统:PdfPig使用左下角为原点的坐标,所以需要按Bottom降序才能得到从上到下的阅读顺序
  • 特殊排版处理:如果PDF存在多列、竖排等复杂排版,需要额外处理X坐标的分组逻辑

内容的提问来源于stack exchange,提问作者Andrus

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.21 11:13:21