You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

iTextSharp提取PDF指定区域文本方向错误问题求助

问题原因与解决建议

问题根源

你遇到的文本顺序错乱问题,核心原因是**LocationTextExtractionStrategy的排序逻辑不适配非标准排版的文本**:

  • 该策略会按文本字符的X/Y坐标位置排序输出,当目标区域文本是垂直排列、反向(从右到左/从下到上)排版时,字符的坐标顺序会和实际阅读顺序相反,导致提取结果乱序。
  • 另外需确认你定义的Rectangle坐标是否符合PDF规则:iTextSharp中PDF坐标原点为页面左下角,如果你的标记是基于左上角原点,会导致区域匹配错误,间接引发文本顺序问题。

解决步骤

1. 替换文本提取策略

改用SimpleTextExtractionStrategy,该策略严格按照PDF内容流中字符出现的顺序提取文本,更适配非标准方向排版:

Dim pageNumber As Integer = 1
Dim rect = New iTextSharp.text.Rectangle(408, 648, 570, 665)
Dim filters As RenderFilter() = {New RegionTextRenderFilter(rect)}
Dim strategy As ITextExtractionStrategy = New FilteredTextRenderListener(
    New SimpleTextExtractionStrategy(), filters)
extractedText = PdfTextExtractor.GetTextFromPage(reader, pageNumber, strategy)

2. 验证并修正区域坐标

检查红色区域的实际坐标:

  • PDF坐标参数为(左下X, 左下Y, 右上X, 右上Y),如果你的标记基于左上角原点,需用页面高度转换Y值:实际Y = 页面高度 - 标记的Y值
  • 可通过reader.GetPageSize(pageNumber).Height获取页面高度,再调整rect的Y参数。

3. 自定义方向感知的提取策略(进阶)

如果上述方法无效,说明文本存在特殊旋转/变换,可自定义策略处理文本方向:

Public Class DirectionAwareTextStrategy
    Inherits LocationTextExtractionStrategy

    Protected Overrides Sub RenderText(ByVal renderInfo As TextRenderInfo)
        Dim textMatrix As Matrix = renderInfo.GetTextMatrix()
        ' 计算文本旋转角度
        Dim rotation As Single = Math.Atan2(textMatrix.Get(Matrix.I21), textMatrix.Get(Matrix.I11)) * (180 / Math.PI)
        Dim text As String = renderInfo.GetText()
        
        ' 针对垂直/反向文本调整字符顺序
        If Math.Abs(rotation - 90) < 1 Or Math.Abs(rotation - 270) < 1 Then
            text = New String(text.Reverse().ToArray())
        End If

        MyBase.RenderText(New TextRenderInfo(renderInfo.GetCharacterRenderInfos(), text, renderInfo.GetTextMatrix(), renderInfo.GetFont(), renderInfo.GetFontSize(), renderInfo.GetWidthOfSpace(), renderInfo.GetSingleSpaceWidth()))
    End Sub
End Class

使用自定义策略:

Dim strategy As ITextExtractionStrategy = New FilteredTextRenderListener(
    New DirectionAwareTextStrategy(), filters)
extractedText = PdfTextExtractor.GetTextFromPage(reader, pageNumber, strategy)

内容的提问来源于stack exchange,提问作者Sven

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.12 19:18:22