如何使用自定义ToUnicode映射提取PDF文本?解决iText7提取问号问题
问题:iText7提取PDF文本输出问号的解决方法?
使用iText7提取PDF文本时,得到的结果全是问号:
???????? ?????????? ? ???????????????????????? ??????????????????? ???????? ???????????????????????????? ...
但在Adobe和Chrome中可正常复制该PDF文本。使用的C#代码如下:
MemoryStream pdfStream = ... pdfStream.Position = 0; var strategy = new LocationTextExtractionStrategy(); var reader = new PdfReader(pdfStream); using var pdfDocument = new PdfDocument(reader); for (int i = 1; i <= pdfDocument.GetNumberOfPages(); ++i) { var page = pdfDocument.GetPage(i); var text = PdfTextExtractor.GetTextFromPage(page, strategy); }
已知该PDF存在字体Unicode映射异常的问题,请问如何用C#实现正确提取?能否通过预处理PDF添加正确映射,或替换为有效映射来解决?
解决方法
1. 自定义iText字体映射修复
针对Unicode映射缺失的问题,可以自定义FontProgram的字符映射逻辑,手动关联字体字形与正确的Unicode字符:
public class CustomFontProgram : FontProgram { private readonly FontProgram _baseFont; private readonly Dictionary<int, int> _glyphToUnicodeMap; public CustomFontProgram(FontProgram baseFont, Dictionary<int, int> glyphToUnicodeMap) { _baseFont = baseFont; _glyphToUnicodeMap = glyphToUnicodeMap; } public override int GetGlyphCode(int unicode) { var glyphEntry = _glyphToUnicodeMap.FirstOrDefault(kv => kv.Value == unicode); return glyphEntry.Key != 0 ? glyphEntry.Key : _baseFont.GetGlyphCode(unicode); } public override int GetUnicode(int glyphCode) { return _glyphToUnicodeMap.TryGetValue(glyphCode, out var unicode) ? unicode : _baseFont.GetUnicode(glyphCode); } public override int GetPdfFontFlags() => _baseFont.GetPdfFontFlags(); public override bool IsFontSpecific(int unicode) => _baseFont.IsFontSpecific(unicode); public override ICollection<int> GetAvailableGlyphs() => _baseFont.GetAvailableGlyphs(); public override int GetKerning(int glyphCode1, int glyphCode2) => _baseFont.GetKerning(glyphCode1, glyphCode2); }
使用时,先通过Chrome/Adobe提取对应页面的正确文本,生成字形码到Unicode的映射表,再替换原字体的映射逻辑:
// 示例映射表,需根据实际PDF内容补充完整 var glyphMap = new Dictionary<int, int> { { 1, 0x41 }, // 字形1对应大写字母A // 补充其他字形与Unicode的对应关系 }; var props = new ReaderProperties(); props.SetFontProgramFactory(new CustomFontProgramFactory(glyphMap)); var reader = new PdfReader(pdfStream, props); using var pdfDocument = new PdfDocument(reader); var strategy = new LocationTextExtractionStrategy(); for (int i = 1; i <= pdfDocument.GetNumberOfPages(); ++i) { var page = pdfDocument.GetPage(i); var text = PdfTextExtractor.GetTextFromPage(page, strategy); }
2. 用OCR工具作为兜底方案
如果自定义映射成本过高,可以结合OCR工具(如Tesseract.NET)对无法提取的页面进行光学字符识别:
// 先用iText尝试提取文本 var text = PdfTextExtractor.GetTextFromPage(page, strategy); if (text.All(c => c == '?')) { // 将PDF页面转为图片 var imageBytes = page.GetImageBytes(); // 使用Tesseract识别图片文本 using var engine = new TesseractEngine("./tessdata", "chi_sim", EngineMode.Default); using var img = Pix.LoadFromMemory(imageBytes); using var ocrPage = engine.Process(img); text = ocrPage.GetText(); }
3. 预处理PDF替换异常字体
可以使用iText的字体替换功能,将有问题的字体替换为系统中带完整Unicode映射的字体:
var writerProps = new WriterProperties(); using var outputStream = new MemoryStream(); using var pdfDocument = new PdfDocument(new PdfReader(pdfStream), new PdfWriter(outputStream, writerProps)); var replacementFont = PdfFontFactory.CreateFont(StandardFonts.HELVETICA); // 遍历每页替换异常字体 for (int i = 1; i <= pdfDocument.GetNumberOfPages(); ++i) { var page = pdfDocument.GetPage(i); var canvasProcessor = new CanvasProcessor(new FontReplacementStrategy(replacementFont)); canvasProcessor.ProcessPageContent(page); } // 保存修复后的PDF,重新读取提取文本 pdfDocument.Close(); outputStream.Position = 0; using var fixedReader = new PdfReader(outputStream); using var fixedDoc = new PdfDocument(fixedReader); // 执行文本提取逻辑...
注:FontReplacementStrategy需继承ICanvasProcessor,实现文本运算符的拦截与字体替换逻辑。
内容的提问来源于stack exchange,提问作者Andrus
相关产品推荐
相关产品推荐

