You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从含空格千位分隔符的PDF提取正确数值(.NET7)

PDF千位分隔符数值提取问题与解决思路

业务场景

业务中需要提取PDF内的Quantity(数量)、Price(单价)、Sum(合计)字段,但不同PDF列布局存在差异,部分PDF以空格作为千位分隔符,示例PDF内容如下:

Description       Price   Quantity          Sum
Soap           1 000.00        2.2     2 200.00
White 3 towel     10.00          2        20.00

现有iText7实现的问题

使用默认LocationTextExtractionStrategy提取文本时,所有内容被单个空格分隔,导致带千位分隔符的数值被拆分,无法直接识别:

提取代码

MemoryStream pdfStream = // 获取PDF文件流
StringBuilder processed = new();
pdfStream.Position = 0;
using var pdfDocument = new PdfDocument(new PdfReader(pdfStream));
var strategy = new LocationTextExtractionStrategy();
for (int i = 1; i <= pdfDocument.GetNumberOfPages(); ++i) {
  var page = pdfDocument.GetPage(i);
  string text = PdfTextExtractor.GetTextFromPage(page, strategy);
  processed.Append(text);
}

提取结果

Soap 1 000.00 2.2 2 200.00
White 3 towel 10.00 2 20.00

由于行内同时存在描述中的普通空格(如White 3 towel)和数值内的千位分隔空格,仅靠纯文本无法区分。

XpdfNet尝试的问题

使用XpdfNet带参数调用时出现文件找不到异常:

System.IO.FileNotFoundException: Could not find file 'C:\myapp\bin\Debug\net7.0\5db7d64c-e1c5-4e1b-b14f-0162ce029c46.txt'.
File name: 'C:\myapp\bin\Debug\net7.0\5db7d64c-e1c5-4e1b-b14f-0162ce029c46.txt'
   at Microsoft.Win32.SafeHandles.SafeFileHandle.CreateFile(String fullPath, FileMode mode, FileAccess access, FileShare share, FileOptions options)
   at Microsoft.Win32.SafeHandles.SafeFileHandle.Open(String fullPath, FileMode mode, FileAccess access, FileShare share, FileOptions options, Int64 preallocationSize, Nullable`1 unixCreateMode)
   at System.IO.Strategies.OSFileStreamStrategy..ctor(String path, FileMode mode, FileAccess access, FileShare share, FileOptions options, Int64 preallocationSize, Nullable`1 unixCreateMode)
   at System.IO.Strategies.FileStreamHelpers.ChooseStrategyCore(String path, FileMode mode, FileAccess access, FileShare share, FileOptions options, Int64 preallocationSize, Nullable`1 unixCreateMode)
   at System.IO.StreamReader.ValidateArgsAndOpenPath(String path, Encoding encoding, Int32 bufferSize)
   at System.IO.File.ReadAllText(String path, Encoding encoding)
   at XpdfNet.XpdfHelper.GetTextResult(XpdfParameter parameter)
   at XpdfNet.XpdfHelper.ToText(String pdfFilePath, String arguments)

无参数调用(string content = pdfHelper.ToText("C:\\a\\test.pdf");)虽运行成功,但提取结果与iText7一致,仍为单个空格分隔的内容。

解决思路:利用文本位置间距区分空格类型

核心是通过获取文本片段的坐标与间距,判断哪些空格是数值内部的千位分隔(间距极小),哪些是列/字段间的分隔(间距较大),具体实现步骤如下:

1. 自定义iText7文本提取策略

继承LocationTextExtractionStrategy,重写RenderText方法,记录每个文本片段的位置信息(X坐标、宽度、Y坐标),用于后续计算间距:

public class SpacingAwareTextExtractionStrategy : LocationTextExtractionStrategy
{
    public List<TextChunk> TextChunks { get; } = new List<TextChunk>();

    public override void RenderText(TextRenderInfo renderInfo)
    {
        base.RenderText(renderInfo);
        var chunk = new TextChunk(renderInfo);
        TextChunks.Add(chunk);
    }

    public class TextChunk
    {
        public string Text { get; }
        public float X { get; }
        public float Width { get; }
        public float Y { get; }

        public TextChunk(TextRenderInfo renderInfo)
        {
            Text = renderInfo.GetText();
            var baseline = renderInfo.GetBaseline();
            X = baseline.GetStartPoint()[0];
            Y = baseline.GetStartPoint()[1];
            Width = renderInfo.GetAdvancedWidth();
        }
    }
}

2. 按行分组并计算间距合并数值

提取所有文本片段后,按Y坐标(行)分组,同一行内按X坐标排序,计算相邻片段的间距,判断是否合并数值:

// 提取文本片段
var strategy = new SpacingAwareTextExtractionStrategy();
PdfTextExtractor.GetTextFromPage(page, strategy);

// 按行分组(Y坐标相近视为同一行)
var rowGroups = strategy.TextChunks
    .GroupBy(c => Math.Round(c.Y, 2)) // 保留两位小数避免精度误差
    .OrderByDescending(g => g.Key); // PDF坐标Y值越大越靠上

foreach (var row in rowGroups)
{
    var sortedChunks = row.OrderBy(c => c.X).ToList();
    StringBuilder lineBuilder = new StringBuilder();
    float previousXEnd = 0;

    foreach (var chunk in sortedChunks)
    {
        if (lineBuilder.Length > 0)
        {
            // 计算当前片段与前一个的间距
            float spacing = chunk.X - previousXEnd;
            // 设定阈值:小于单个字符宽度的1/2视为数值内部空格
            float charWidth = chunk.Width / chunk.Text.Length;
            var prevChunk = sortedChunks[sortedChunks.IndexOf(chunk)-1];
            
            if (spacing < charWidth / 2 && IsNumericPart(chunk.Text) && IsNumericPart(prevChunk.Text))
            {
                // 合并数值内部空格
                lineBuilder.Append(chunk.Text);
            }
            else
            {
                // 字段间空格,保留分隔
                lineBuilder.Append(" ").Append(chunk.Text);
            }
        }
        else
        {
            lineBuilder.Append(chunk.Text);
        }
        previousXEnd = chunk.X + chunk.Width;
    }
    processed.AppendLine(lineBuilder.ToString());
}

// 辅助方法:判断是否为数值片段(数字、小数点)
private bool IsNumericPart(string text)
{
    return text.All(c => char.IsDigit(c) || c == '.');
}

3. 结合列位置优化(可选)

如果PDF列布局相对固定,可以先通过表头确定Price、Quantity、Sum列的X坐标范围,仅在这些范围内处理数值合并,避免误合并描述中的数字(如White 3 towel里的3):

// 从表头获取列的X范围
var headerChunks = rowGroups.First().OrderBy(c => c.X).ToList();
float priceColumnStart = headerChunks.First(c => c.Text == "Price").X;
float quantityColumnStart = headerChunks.First(c => c.Text == "Quantity").X;
float sumColumnStart = headerChunks.First(c => c.Text == "Sum").X;

// 合并时增加列范围判断
if (spacing < charWidth / 2 && IsNumericPart(chunk.Text) && IsNumericPart(prevChunk.Text) 
    && (chunk.X >= priceColumnStart || chunk.X >= quantityColumnStart || chunk.X >= sumColumnStart))
{
    // 合并数值
}

关键说明

  • 间距阈值需根据实际PDF的字体大小调整,建议通过测试样本校准;
  • 对于布局差异大的PDF,可结合正则匹配数值模式进一步优化;
  • XpdfNet的文件找不到异常需检查工具依赖的本地文件是否存在,或改用官方命令行工具直接调用。

内容的提问来源于stack exchange,提问作者Andrus

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.23 09:58:10