基于iTextSharp的PDF教材题目提取与结构化处理技术咨询
开发需求与问题
我要开发一款程序,能按主题和难度从教材PDF中随机提取题目。目前已实现两个功能:按页码输出页面文本、按页码与文本提取文本块坐标。但还有几个核心难题待解决:
- 如何让程序识别题目所属的习题章节
- 如何提取包含上下标、加减以外的数学符号等复杂字符的题目文本
- 若题目包含图表,如何将题目(含图表)截取为图片
典型问题
提取的题目里,下标会被拆成独立文本块,教材图表上的文本(比如P₁、P₂)也会被识别成独立块,导致提取结果混乱,示例输出如下:
- A triangle has vertices at points in the Argand diagram which represent the complex numbers z₁, z₂ and z₃.
y
z₂-z₁
π π
If =cos +isin , show that the triangle is equilateral.
3 3
P
z₃-z₁
2
P₃- In the diagram on the right, the points P₁, P₂ and P₃ represent
the complex numbers z₁, z₂ and z₃ respectively. If z₂/z₁ = z₃/z₂, show
that OP₂ bisects P₁OP₃.
现有实现代码
我用iTextSharp实现现有功能,代码如下:
TextExtractionStrategy.vb
Imports System Imports System.Collections.Generic Imports iTextSharp.text.pdf.parser Namespace TextExtractionStrategy Public Class LocationTextExtractionStrategyWithPosition Inherits LocationTextExtractionStrategy Private ReadOnly locationalResult As List(Of TextChunk) = New List(Of TextChunk)() Private ReadOnly tclStrat As ITextChunkLocationStrategy Public Sub New() 'constructors' Me.New(New TextChunkLocationStrategyDefaultImp()) End Sub Public Sub New(ByVal strat As ITextChunkLocationStrategy) tclStrat = strat End Sub Private Function StartsWithSpace(ByVal str As String) As Boolean 'Logical Operators to check for spaces' If str.Length = 0 Then Return False Return str(0) = " "c End Function Private Function EndsWithSpace(ByVal str As String) As Boolean If str.Length = 0 Then Return False Return str(str.Length - 1) = " "c End Function Private Function filterTextChunks(ByVal textChunks As List(Of TextChunk), ByVal filter As ITextChunkFilter) As List(Of TextChunk) If filter Is Nothing Then 'does nothing if no filters are applied' Return textChunks End If Dim filtered = New List(Of TextChunk)() For Each textChunk In textChunks 'checks chunks for if they apply to filter' If filter.Accept(textChunk) Then filtered.Add(textChunk) End If Next Return filtered End Function Public Overrides Sub RenderText(ByVal renderInfo As TextRenderInfo) Dim segment As LineSegment = renderInfo.GetBaseline() If renderInfo.GetRise() <> 0 Then Dim riseOffsetTransform As Matrix = New Matrix(0, -renderInfo.GetRise()) segment = segment.TransformBy(riseOffsetTransform) End If Dim tc As TextChunk = New TextChunk(renderInfo.GetText(), tclStrat.CreateLocation(renderInfo, segment)) locationalResult.Add(tc) End Sub Public Function GetLocations() As IList(Of TextLocation) Dim filteredTextChunks = filterTextChunks(locationalResult, Nothing) filteredTextChunks.Sort() 'sorts text chunks' Dim lastChunk As TextChunk = Nothing Dim textLocations = New List(Of TextLocation)() For Each chunk In filteredTextChunks If lastChunk Is Nothing Then 'add the first chunk' textLocations.Add(New TextLocation With { .Text = chunk.Text, .X = chunk.Location.StartLocation(0), .Y = chunk.Location.StartLocation(1) }) Else If chunk.SameLine(lastChunk) Then 'if the chunk is on the same line as the previous chunk' Dim text = "" 'clear text' If IsChunkAtWordBoundary(chunk, lastChunk) AndAlso Not StartsWithSpace(chunk.Text) AndAlso Not EndsWithSpace(lastChunk.Text) Then text += " "c 'add a space if it doesnt already have one where it needs to be' text += chunk.Text 'add text to space' textLocations(textLocations.Count - 1).Text += text 'add text to the previous chunk' Else 'otherwise the chunk is on a new line, so it can be added as a brand new chunk' textLocations.Add(New TextLocation With { .Text = chunk.Text, .X = chunk.Location.StartLocation(0), .Y = chunk.Location.StartLocation(1) }) End If End If lastChunk = chunk Next Return textLocations End Function End Class Public Class TextLocation 'Custom class containing text and its xy coords' Public Property X As Single Public Property Y As Single Public Property Text As String End Class End Namespace
Program.vb
Imports System Imports System.Collections.Generic Imports iTextSharp.text.pdf.parser Imports iTextSharp.text.pdf Imports System.Text Namespace TextExtractionStrategy Module Program Dim PDFLocation As String = "C:\math.pdf" Sub Main(args As String()) Dim codepages = CodePagesEncodingProvider.Instance Encoding.RegisterProvider(codepages) ReadText() End Sub Function GetTextCoord(page As Integer, searchText As String) As List(Of TextLocation) Dim searchResult As List(Of TextLocation) = New List(Of TextLocation) Using reader As PdfReader = New PdfReader(PDFLocation) Dim parser = New PdfReaderContentParser(reader) Dim strategy = parser.ProcessContent(page, New LocationTextExtractionStrategyWithPosition()) Dim res = strategy.GetLocations() reader.Close() Dim temp As TextLocation() = res.ToArray For i = 0 To temp.Length - 1 If temp(i).Text.Contains(searchText) Then searchResult.Add(temp(i)) End If Next End Using Return searchResult End Function Sub ReadText() Using reader As PdfReader = New PdfReader("C:\math.pdf") Dim text = PdfTextExtractor.GetTextFromPage(reader, 46) reader.Close() Console.WriteLine(text) End Using End Sub End Module End Namespace
依赖与测试文件
- 依赖包:iTextSharp、System.Text.Encoding.CodePages
- 测试用PDF教材需放在
C:\math.pdf路径下
解决思路
1. 识别习题章节
- 匹配章节标题格式:通过正则表达式匹配统一格式的章节标题(如"Exercise X.X"、"习题X-X"),结合文本块的垂直坐标范围确定章节的起始与结束页码。
- 利用PDF书签结构:通过
PdfReader.GetOutlines()获取文档的层级书签,直接映射到对应章节的页码区间,再在区间内提取题目。
2. 修复上下标与复杂数学符号提取问题
- 上下标合并:在
RenderText方法中,通过renderInfo.GetRise()判断文本块是上标(rise为正)还是下标(rise为负),记录其与基准文本块的关联;合并文本时,将上下标文本直接拼接在对应基准文本的对应位置,而非作为独立行处理。 - 复杂符号处理:确保编码支持已配置(已注册
CodePagesEncodingProvider),检查PDF字体是否完整嵌入;对于无法直接提取的符号,可通过renderInfo.GetFont()获取字体编码映射还原,或结合OCR工具辅助提取。
3. 含图表的题目截图
- 确定题目区域:先通过文本提取确定题号(如"16.")和下一题题号(如"17.")的坐标,扩展范围包含周边图表(可通过文本块间距、PDF图像元素位置判断)。
- 截取页面区域:使用
PdfStamper创建仅包含目标区域的PDF子文档,再结合Ghostscript或System.Drawing将该子文档页面转换为图片保存。
内容的提问来源于stack exchange,提问作者LBloxo
相关产品推荐
相关产品推荐

