如何用iText在PowerShell中读取PDF文本行的颜色?
解决PowerShell+iTextSharp提取PDF带颜色文本的问题
默认的PdfTextExtractor.GetTextFromPage只能提取纯文本,无法获取文本颜色。要实现带颜色的文本提取,需要自定义文本提取策略,继承LocationTextExtractionStrategy并捕获文本对应的颜色信息。
核心实现步骤
- 定义自定义策略类,重写
RenderText方法,记录每个文本片段的内容和颜色 - 使用自定义策略提取每页的文本及颜色
- 过滤出红色或蓝色的文本行
完整PowerShell代码
Add-Type -Path "C:\Users\Ion\Documents\App\Scripts\itextsharp.dll" # 定义自定义文本提取策略,捕获文本和颜色 $colorTextStrategy = @" using System.Collections.Generic; using iTextSharp.text.pdf; using iTextSharp.text.pdf.parser; using iTextSharp.text; public class ColorTextExtractionStrategy : LocationTextExtractionStrategy { public List<TextWithColor> TextFragments { get; } = new List<TextWithColor>(); public override void RenderText(TextRenderInfo renderInfo) { base.RenderText(renderInfo); // 获取文本填充色(大多数PDF文本用填充色) BaseColor fillColor = renderInfo.GetFillColor(); // 若填充色为空,尝试获取描边色 if (fillColor == null) { fillColor = renderInfo.GetStrokeColor(); } TextFragments.Add(new TextWithColor { Text = renderInfo.GetText(), Color = fillColor, YPosition = renderInfo.GetBaseline().GetStartPoint()[1] }); } } public class TextWithColor { public string Text { get; set; } public BaseColor Color { get; set; } public float YPosition { get; set; } } "@ Add-Type -TypeDefinition $colorTextStrategy -ReferencedAssemblies "C:\Users\Ion\Documents\App\Scripts\itextsharp.dll" # PDF文件路径 $filePath = "C:\Users\Scripts\Datos\ADMINISTRATIVO-AEPSA-SERV.-CENTRALES-modificado.pdf" $outputPath = "C:\Users\Scripts\Datos\datos_with_color.txt" $pdf = New-Object iTextSharp.text.pdf.PdfReader -ArgumentList $filePath $export = "" foreach ($page in 1..$pdf.NumberOfPages) { $strategy = New-Object ColorTextExtractionStrategy [iTextSharp.text.pdf.parser.PdfTextExtractor]::GetTextFromPage($pdf, $page, $strategy) # 按Y坐标分组(同一行的文本Y坐标接近),合并成完整行并判断颜色 $groupedByLine = $strategy.TextFragments | Group-Object -Property { [math]::Round($_.YPosition, 1) } | Sort-Object -Property Name -Descending foreach ($lineGroup in $groupedByLine) { $lineText = ($lineGroup.Group | Sort-Object -Property { $_.Text.Length } -Descending).Text -join "" $lineColor = $lineGroup.Group[0].Color # 判断是否为红色或蓝色(RGB值判断,可根据实际PDF颜色调整) $isRed = $lineColor -ne $null -and $lineColor.R -eq 255 -and $lineColor.G -eq 0 -and $lineColor.B -eq 0 $isBlue = $lineColor -ne $null -and $lineColor.R -eq 0 -and $lineColor.G -eq 0 -and $lineColor.B -eq 255 if ($isRed -or $isBlue) { $export += "[彩色行: $($isRed ? '红色' : '蓝色')] $lineText`r`n" } else { $export += "$lineText`r`n" } } } $pdf.Close() $export | Out-File -FilePath $outputPath -Encoding UTF8
代码说明
- 自定义策略类:
ColorTextExtractionStrategy继承自LocationTextExtractionStrategy,在RenderText中捕获每个文本片段的填充色(优先)或描边色,同时记录Y坐标用于合并行。 - 行合并逻辑:同一行的文本Y坐标高度接近,通过对Y坐标取整分组,合并成完整的文本行。
- 颜色判断:通过RGB值判断纯红色/纯蓝色,若PDF使用非标准RGB值,可修改为近似匹配逻辑(如判断R值远高于G/B即为红色)。
其他自动化方案
- PdfSharp:另一款.NET PDF处理库,支持类似的自定义文本提取逻辑,语法与iTextSharp接近。
- Python PyMuPDF(fitz):若可切换语言,PyMuPDF提取带颜色文本更简洁,示例代码:
import fitz doc = fitz.open("your_pdf_path.pdf") output = "" for page in doc: blocks = page.get_text("dict")["blocks"] for block in blocks: if "lines" in block: for line in block["lines"]: for span in line["spans"]: color = span["color"] # 转换fitz整数颜色为RGB r = (color >> 16) & 0xff g = (color >> 8) & 0xff b = color & 0xff is_red = r == 255 and g == 0 and b == 0 is_blue = r == 0 and g == 0 and b == 255 if is_red or is_blue: output += f"[彩色行: {'红色' if is_red else '蓝色'}] {span['text']}\n" else: output += f"{span['text']}\n" with open("output.txt", "w", encoding="utf-8") as f: f.write(output)
内容的提问来源于stack exchange,提问作者user3187095
相关产品推荐
相关产品推荐

