You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用自定义ToUnicode映射提取PDF文本?解决iText7提取问号问题

问题:iText7提取PDF文本输出问号的解决方法?

使用iText7提取PDF文本时,得到的结果全是问号:

???????? ??????????
?
???????????????????????? ???????????????????
???????? ????????????????????????????
...

但在Adobe和Chrome中可正常复制该PDF文本。使用的C#代码如下:

MemoryStream pdfStream = ...
pdfStream.Position = 0;
var strategy = new LocationTextExtractionStrategy();
var reader = new PdfReader(pdfStream);
using var pdfDocument = new PdfDocument(reader);
for (int i = 1; i <= pdfDocument.GetNumberOfPages(); ++i)
{
    var page = pdfDocument.GetPage(i);
    var text = PdfTextExtractor.GetTextFromPage(page, strategy);
}

已知该PDF存在字体Unicode映射异常的问题,请问如何用C#实现正确提取?能否通过预处理PDF添加正确映射,或替换为有效映射来解决?


解决方法

1. 自定义iText字体映射修复

针对Unicode映射缺失的问题,可以自定义FontProgram的字符映射逻辑,手动关联字体字形与正确的Unicode字符:

public class CustomFontProgram : FontProgram
{
    private readonly FontProgram _baseFont;
    private readonly Dictionary<int, int> _glyphToUnicodeMap;

    public CustomFontProgram(FontProgram baseFont, Dictionary<int, int> glyphToUnicodeMap)
    {
        _baseFont = baseFont;
        _glyphToUnicodeMap = glyphToUnicodeMap;
    }

    public override int GetGlyphCode(int unicode)
    {
        var glyphEntry = _glyphToUnicodeMap.FirstOrDefault(kv => kv.Value == unicode);
        return glyphEntry.Key != 0 ? glyphEntry.Key : _baseFont.GetGlyphCode(unicode);
    }

    public override int GetUnicode(int glyphCode)
    {
        return _glyphToUnicodeMap.TryGetValue(glyphCode, out var unicode) ? unicode : _baseFont.GetUnicode(glyphCode);
    }

    public override int GetPdfFontFlags() => _baseFont.GetPdfFontFlags();
    public override bool IsFontSpecific(int unicode) => _baseFont.IsFontSpecific(unicode);
    public override ICollection<int> GetAvailableGlyphs() => _baseFont.GetAvailableGlyphs();
    public override int GetKerning(int glyphCode1, int glyphCode2) => _baseFont.GetKerning(glyphCode1, glyphCode2);
}

使用时,先通过Chrome/Adobe提取对应页面的正确文本,生成字形码到Unicode的映射表,再替换原字体的映射逻辑:

// 示例映射表,需根据实际PDF内容补充完整
var glyphMap = new Dictionary<int, int>
{
    { 1, 0x41 }, // 字形1对应大写字母A
    // 补充其他字形与Unicode的对应关系
};

var props = new ReaderProperties();
props.SetFontProgramFactory(new CustomFontProgramFactory(glyphMap));
var reader = new PdfReader(pdfStream, props);

using var pdfDocument = new PdfDocument(reader);
var strategy = new LocationTextExtractionStrategy();
for (int i = 1; i <= pdfDocument.GetNumberOfPages(); ++i)
{
    var page = pdfDocument.GetPage(i);
    var text = PdfTextExtractor.GetTextFromPage(page, strategy);
}

2. 用OCR工具作为兜底方案

如果自定义映射成本过高,可以结合OCR工具(如Tesseract.NET)对无法提取的页面进行光学字符识别:

// 先用iText尝试提取文本
var text = PdfTextExtractor.GetTextFromPage(page, strategy);
if (text.All(c => c == '?'))
{
    // 将PDF页面转为图片
    var imageBytes = page.GetImageBytes();
    // 使用Tesseract识别图片文本
    using var engine = new TesseractEngine("./tessdata", "chi_sim", EngineMode.Default);
    using var img = Pix.LoadFromMemory(imageBytes);
    using var ocrPage = engine.Process(img);
    text = ocrPage.GetText();
}

3. 预处理PDF替换异常字体

可以使用iText的字体替换功能,将有问题的字体替换为系统中带完整Unicode映射的字体:

var writerProps = new WriterProperties();
using var outputStream = new MemoryStream();
using var pdfDocument = new PdfDocument(new PdfReader(pdfStream), new PdfWriter(outputStream, writerProps));
var replacementFont = PdfFontFactory.CreateFont(StandardFonts.HELVETICA);

// 遍历每页替换异常字体
for (int i = 1; i <= pdfDocument.GetNumberOfPages(); ++i)
{
    var page = pdfDocument.GetPage(i);
    var canvasProcessor = new CanvasProcessor(new FontReplacementStrategy(replacementFont));
    canvasProcessor.ProcessPageContent(page);
}

// 保存修复后的PDF,重新读取提取文本
pdfDocument.Close();
outputStream.Position = 0;
using var fixedReader = new PdfReader(outputStream);
using var fixedDoc = new PdfDocument(fixedReader);
// 执行文本提取逻辑...

注:FontReplacementStrategy需继承ICanvasProcessor,实现文本运算符的拦截与字体替换逻辑。


内容的提问来源于stack exchange,提问作者Andrus

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.22 22:55:11