如何从含空格千位分隔符的PDF提取正确数值(.NET7)
PDF千位分隔符数值提取问题与解决思路
业务场景
业务中需要提取PDF内的Quantity(数量)、Price(单价)、Sum(合计)字段,但不同PDF列布局存在差异,部分PDF以空格作为千位分隔符,示例PDF内容如下:
Description Price Quantity Sum Soap 1 000.00 2.2 2 200.00 White 3 towel 10.00 2 20.00
现有iText7实现的问题
使用默认LocationTextExtractionStrategy提取文本时,所有内容被单个空格分隔,导致带千位分隔符的数值被拆分,无法直接识别:
提取代码
MemoryStream pdfStream = // 获取PDF文件流 StringBuilder processed = new(); pdfStream.Position = 0; using var pdfDocument = new PdfDocument(new PdfReader(pdfStream)); var strategy = new LocationTextExtractionStrategy(); for (int i = 1; i <= pdfDocument.GetNumberOfPages(); ++i) { var page = pdfDocument.GetPage(i); string text = PdfTextExtractor.GetTextFromPage(page, strategy); processed.Append(text); }
提取结果
Soap 1 000.00 2.2 2 200.00 White 3 towel 10.00 2 20.00
由于行内同时存在描述中的普通空格(如White 3 towel)和数值内的千位分隔空格,仅靠纯文本无法区分。
XpdfNet尝试的问题
使用XpdfNet带参数调用时出现文件找不到异常:
System.IO.FileNotFoundException: Could not find file 'C:\myapp\bin\Debug\net7.0\5db7d64c-e1c5-4e1b-b14f-0162ce029c46.txt'. File name: 'C:\myapp\bin\Debug\net7.0\5db7d64c-e1c5-4e1b-b14f-0162ce029c46.txt' at Microsoft.Win32.SafeHandles.SafeFileHandle.CreateFile(String fullPath, FileMode mode, FileAccess access, FileShare share, FileOptions options) at Microsoft.Win32.SafeHandles.SafeFileHandle.Open(String fullPath, FileMode mode, FileAccess access, FileShare share, FileOptions options, Int64 preallocationSize, Nullable`1 unixCreateMode) at System.IO.Strategies.OSFileStreamStrategy..ctor(String path, FileMode mode, FileAccess access, FileShare share, FileOptions options, Int64 preallocationSize, Nullable`1 unixCreateMode) at System.IO.Strategies.FileStreamHelpers.ChooseStrategyCore(String path, FileMode mode, FileAccess access, FileShare share, FileOptions options, Int64 preallocationSize, Nullable`1 unixCreateMode) at System.IO.StreamReader.ValidateArgsAndOpenPath(String path, Encoding encoding, Int32 bufferSize) at System.IO.File.ReadAllText(String path, Encoding encoding) at XpdfNet.XpdfHelper.GetTextResult(XpdfParameter parameter) at XpdfNet.XpdfHelper.ToText(String pdfFilePath, String arguments)
无参数调用(string content = pdfHelper.ToText("C:\\a\\test.pdf");)虽运行成功,但提取结果与iText7一致,仍为单个空格分隔的内容。
解决思路:利用文本位置间距区分空格类型
核心是通过获取文本片段的坐标与间距,判断哪些空格是数值内部的千位分隔(间距极小),哪些是列/字段间的分隔(间距较大),具体实现步骤如下:
1. 自定义iText7文本提取策略
继承LocationTextExtractionStrategy,重写RenderText方法,记录每个文本片段的位置信息(X坐标、宽度、Y坐标),用于后续计算间距:
public class SpacingAwareTextExtractionStrategy : LocationTextExtractionStrategy { public List<TextChunk> TextChunks { get; } = new List<TextChunk>(); public override void RenderText(TextRenderInfo renderInfo) { base.RenderText(renderInfo); var chunk = new TextChunk(renderInfo); TextChunks.Add(chunk); } public class TextChunk { public string Text { get; } public float X { get; } public float Width { get; } public float Y { get; } public TextChunk(TextRenderInfo renderInfo) { Text = renderInfo.GetText(); var baseline = renderInfo.GetBaseline(); X = baseline.GetStartPoint()[0]; Y = baseline.GetStartPoint()[1]; Width = renderInfo.GetAdvancedWidth(); } } }
2. 按行分组并计算间距合并数值
提取所有文本片段后,按Y坐标(行)分组,同一行内按X坐标排序,计算相邻片段的间距,判断是否合并数值:
// 提取文本片段 var strategy = new SpacingAwareTextExtractionStrategy(); PdfTextExtractor.GetTextFromPage(page, strategy); // 按行分组(Y坐标相近视为同一行) var rowGroups = strategy.TextChunks .GroupBy(c => Math.Round(c.Y, 2)) // 保留两位小数避免精度误差 .OrderByDescending(g => g.Key); // PDF坐标Y值越大越靠上 foreach (var row in rowGroups) { var sortedChunks = row.OrderBy(c => c.X).ToList(); StringBuilder lineBuilder = new StringBuilder(); float previousXEnd = 0; foreach (var chunk in sortedChunks) { if (lineBuilder.Length > 0) { // 计算当前片段与前一个的间距 float spacing = chunk.X - previousXEnd; // 设定阈值:小于单个字符宽度的1/2视为数值内部空格 float charWidth = chunk.Width / chunk.Text.Length; var prevChunk = sortedChunks[sortedChunks.IndexOf(chunk)-1]; if (spacing < charWidth / 2 && IsNumericPart(chunk.Text) && IsNumericPart(prevChunk.Text)) { // 合并数值内部空格 lineBuilder.Append(chunk.Text); } else { // 字段间空格,保留分隔 lineBuilder.Append(" ").Append(chunk.Text); } } else { lineBuilder.Append(chunk.Text); } previousXEnd = chunk.X + chunk.Width; } processed.AppendLine(lineBuilder.ToString()); } // 辅助方法:判断是否为数值片段(数字、小数点) private bool IsNumericPart(string text) { return text.All(c => char.IsDigit(c) || c == '.'); }
3. 结合列位置优化(可选)
如果PDF列布局相对固定,可以先通过表头确定Price、Quantity、Sum列的X坐标范围,仅在这些范围内处理数值合并,避免误合并描述中的数字(如White 3 towel里的3):
// 从表头获取列的X范围 var headerChunks = rowGroups.First().OrderBy(c => c.X).ToList(); float priceColumnStart = headerChunks.First(c => c.Text == "Price").X; float quantityColumnStart = headerChunks.First(c => c.Text == "Quantity").X; float sumColumnStart = headerChunks.First(c => c.Text == "Sum").X; // 合并时增加列范围判断 if (spacing < charWidth / 2 && IsNumericPart(chunk.Text) && IsNumericPart(prevChunk.Text) && (chunk.X >= priceColumnStart || chunk.X >= quantityColumnStart || chunk.X >= sumColumnStart)) { // 合并数值 }
关键说明
- 间距阈值需根据实际PDF的字体大小调整,建议通过测试样本校准;
- 对于布局差异大的PDF,可结合正则匹配数值模式进一步优化;
- XpdfNet的文件找不到异常需检查工具依赖的本地文件是否存在,或改用官方命令行工具直接调用。
内容的提问来源于stack exchange,提问作者Andrus
相关产品推荐
相关产品推荐

