You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用iText7提取PDF文本时回车换行丢失的问题

PDF文本提取丢失连续换行问题

问题代码

public List<string> ExtractTextFromPdfA(string pdfPath)
{
    var extractedText = new List<string>();
    try
    {
        using (PdfReader pdfReader = new PdfReader(pdfPath))
        using (PdfDocument pdfDocument = new PdfDocument(pdfReader))
        {
            for (int i = 1; i <= pdfDocument.GetNumberOfPages(); i++)
            {
                //ITextExtractionStrategy strategy = new LocationTextExtractionStrategy();
                ITextExtractionStrategy strategy = new SimpleTextExtractionStrategy();
                string pageText = PdfTextExtractor.GetTextFromPage(pdfDocument.GetPage(i), strategy);
                extractedText.Add(pageText);
            }
            return extractedText;
        }
    }
    catch (Exception ex)
    {
        //_logger.LogError(ex, "Error extracting text from pdf");
        throw;
    }
}

问题描述

上述代码使用iText的SimpleTextExtractionStrategy提取PDF文本时,原内容中的连续换行(\n\n)或Windows格式换行(\r\n)均被压缩为单个\n,导致原始文本的段落分隔等格式丢失。

问题原因

SimpleTextExtractionStrategy的设计逻辑是合并冗余的空白和换行,只保留最基础的文本换行,目的是输出连续易读的文本,但会忽略PDF中原本的段落级换行分隔。

解决方案

方案1:自定义文本提取策略(推荐)

通过继承LocationTextExtractionStrategy,重写文本渲染逻辑,根据文本块的垂直间距判断是否需要添加额外换行,从而保留原始段落分隔:

public class PreserveLineBreaksStrategy : LocationTextExtractionStrategy
{
    private float _lastTextYPosition;
    // 垂直间距阈值,可根据目标PDF的排版调整
    private const float ParagraphBreakThreshold = 6f;

    public override void RenderText(TextRenderInfo renderInfo)
    {
        base.RenderText(renderInfo);
        // 获取当前文本块的底部Y坐标
        float currentY = renderInfo.GetDescentLine().GetStartPoint()[1];
        
        // 对比上一个文本块的Y坐标,间距超过阈值则插入额外换行
        if (_lastTextYPosition != 0 && _lastTextYPosition - currentY > ParagraphBreakThreshold)
        {
            AppendText("\n");
        }
        
        _lastTextYPosition = currentY;
    }
}

修改原代码中的策略实例化:

ITextExtractionStrategy strategy = new PreserveLineBreaksStrategy();

方案2:提取后文本补全换行(简易粗暴)

如果不需要精确识别段落,只是需要将单个换行还原为连续换行,可以在提取文本后直接替换:

string pageText = PdfTextExtractor.GetTextFromPage(pdfDocument.GetPage(i), strategy);
// 将单个换行替换为连续换行,注意可能会影响正常行间距
pageText = pageText.Replace("\n", "\n\n");
extractedText.Add(pageText);

注意事项

  • 不同PDF的排版差异较大,ParagraphBreakThreshold的数值需要根据目标PDF实际调整,行间距大的PDF需要调高阈值。
  • 如果是扫描版PDF,iText无法直接提取文本,需要先通过OCR工具识别文本,再处理换行问题。

内容的提问来源于stack exchange,提问作者user29888433

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.14 03:06:01