You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用iText在PowerShell中读取PDF文本行的颜色?

解决PowerShell+iTextSharp提取PDF带颜色文本的问题

默认的PdfTextExtractor.GetTextFromPage只能提取纯文本,无法获取文本颜色。要实现带颜色的文本提取,需要自定义文本提取策略,继承LocationTextExtractionStrategy并捕获文本对应的颜色信息。

核心实现步骤

  • 定义自定义策略类,重写RenderText方法,记录每个文本片段的内容和颜色
  • 使用自定义策略提取每页的文本及颜色
  • 过滤出红色或蓝色的文本行

完整PowerShell代码

Add-Type -Path "C:\Users\Ion\Documents\App\Scripts\itextsharp.dll"

# 定义自定义文本提取策略,捕获文本和颜色
$colorTextStrategy = @"
using System.Collections.Generic;
using iTextSharp.text.pdf;
using iTextSharp.text.pdf.parser;
using iTextSharp.text;

public class ColorTextExtractionStrategy : LocationTextExtractionStrategy
{
    public List<TextWithColor> TextFragments { get; } = new List<TextWithColor>();

    public override void RenderText(TextRenderInfo renderInfo)
    {
        base.RenderText(renderInfo);
        // 获取文本填充色(大多数PDF文本用填充色)
        BaseColor fillColor = renderInfo.GetFillColor();
        // 若填充色为空,尝试获取描边色
        if (fillColor == null)
        {
            fillColor = renderInfo.GetStrokeColor();
        }
        TextFragments.Add(new TextWithColor
        {
            Text = renderInfo.GetText(),
            Color = fillColor,
            YPosition = renderInfo.GetBaseline().GetStartPoint()[1]
        });
    }
}

public class TextWithColor
{
    public string Text { get; set; }
    public BaseColor Color { get; set; }
    public float YPosition { get; set; }
}
"@

Add-Type -TypeDefinition $colorTextStrategy -ReferencedAssemblies "C:\Users\Ion\Documents\App\Scripts\itextsharp.dll"

# PDF文件路径
$filePath = "C:\Users\Scripts\Datos\ADMINISTRATIVO-AEPSA-SERV.-CENTRALES-modificado.pdf"
$outputPath = "C:\Users\Scripts\Datos\datos_with_color.txt"

$pdf = New-Object iTextSharp.text.pdf.PdfReader -ArgumentList $filePath
$export = ""

foreach ($page in 1..$pdf.NumberOfPages) {
    $strategy = New-Object ColorTextExtractionStrategy
    [iTextSharp.text.pdf.parser.PdfTextExtractor]::GetTextFromPage($pdf, $page, $strategy)
    
    # 按Y坐标分组(同一行的文本Y坐标接近),合并成完整行并判断颜色
    $groupedByLine = $strategy.TextFragments | Group-Object -Property { [math]::Round($_.YPosition, 1) } | Sort-Object -Property Name -Descending
    
    foreach ($lineGroup in $groupedByLine) {
        $lineText = ($lineGroup.Group | Sort-Object -Property { $_.Text.Length } -Descending).Text -join ""
        $lineColor = $lineGroup.Group[0].Color
        
        # 判断是否为红色或蓝色(RGB值判断,可根据实际PDF颜色调整)
        $isRed = $lineColor -ne $null -and $lineColor.R -eq 255 -and $lineColor.G -eq 0 -and $lineColor.B -eq 0
        $isBlue = $lineColor -ne $null -and $lineColor.R -eq 0 -and $lineColor.G -eq 0 -and $lineColor.B -eq 255
        
        if ($isRed -or $isBlue) {
            $export += "[彩色行: $($isRed ? '红色' : '蓝色')] $lineText`r`n"
        } else {
            $export += "$lineText`r`n"
        }
    }
}

$pdf.Close()
$export | Out-File -FilePath $outputPath -Encoding UTF8

代码说明

  • 自定义策略类:ColorTextExtractionStrategy继承自LocationTextExtractionStrategy,在RenderText中捕获每个文本片段的填充色(优先)或描边色,同时记录Y坐标用于合并行。
  • 行合并逻辑:同一行的文本Y坐标高度接近,通过对Y坐标取整分组,合并成完整的文本行。
  • 颜色判断:通过RGB值判断纯红色/纯蓝色,若PDF使用非标准RGB值,可修改为近似匹配逻辑(如判断R值远高于G/B即为红色)。

其他自动化方案

  • PdfSharp:另一款.NET PDF处理库,支持类似的自定义文本提取逻辑,语法与iTextSharp接近。
  • Python PyMuPDF(fitz):若可切换语言,PyMuPDF提取带颜色文本更简洁,示例代码:
import fitz

doc = fitz.open("your_pdf_path.pdf")
output = ""
for page in doc:
    blocks = page.get_text("dict")["blocks"]
    for block in blocks:
        if "lines" in block:
            for line in block["lines"]:
                for span in line["spans"]:
                    color = span["color"]
                    # 转换fitz整数颜色为RGB
                    r = (color >> 16) & 0xff
                    g = (color >> 8) & 0xff
                    b = color & 0xff
                    is_red = r == 255 and g == 0 and b == 0
                    is_blue = r == 0 and g == 0 and b == 255
                    if is_red or is_blue:
                        output += f"[彩色行: {'红色' if is_red else '蓝色'}] {span['text']}\n"
                    else:
                        output += f"{span['text']}\n"
with open("output.txt", "w", encoding="utf-8") as f:
    f.write(output)

内容的提问来源于stack exchange,提问作者user3187095

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.06 22:05:27