如何使用iText7获取PDF中Type 3字体的实际名称?
获取iText7中Type 3字体的名称
Great question! I’ve dealt with this exact quirk of Type 3 fonts in iText7 before—they don’t expose their names through the standard GetFontNames() method because they’re embedded as custom, user-defined fonts rather than pre-defined font programs. Acrobat Reader pulls the name directly from the underlying PDF dictionary, and we can do the same in C#.
问题原因
Type 3字体在iText7中由Type3Font实例表示,它们的FontProgram(即Type3FontProgram)不会填充GetFontNames()所需的字体名称数据。相反,字体名称存储在PDF的BaseFont字典条目中的Name键下。
修改后的解决方案代码
下面是更新后的FontReader类代码,它可以同时捕获标准字体和Type 3字体的名称:
private class FontReader : IEventListener { public ICollection<string> Fonts { get; } public FontReader() { Fonts = new List<string>(); } public void EventOccurred(IEventData data, EventType type) { if (!(data is TextRenderInfo textRenderInfo)) return; var font = textRenderInfo.GetFont(); string fontName = null; // 处理标准字体类型(Type 1、TrueType等) var fontProgram = font.GetFontProgram(); if (!(fontProgram is Type3FontProgram)) { fontName = fontProgram.GetFontNames().GetFontName(); } // 通过直接读取PDF字典处理Type 3字体 else if (font is Type3Font type3Font) { var fontDict = type3Font.GetPdfObject() as PdfDictionary; if (fontDict != null) { var baseFont = fontDict.GetAsDictionary(PdfName.BaseFont); if (baseFont != null) { var nameEntry = baseFont.GetAsString(PdfName.Name); fontName = nameEntry?.GetValue(); } } } // 如果名称有效且未在集合中,则添加 if (!string.IsNullOrEmpty(fontName) && !Fonts.Contains(fontName)) { Fonts.Add(fontName); } } public ICollection<EventType> GetSupportedEvents() { return new HashSet<EventType> {EventType.RENDER_TEXT}; } }
关键步骤说明
- 判断Type 3字体:首先验证字体是否为
Type3Font实例,这意味着我们需要访问底层的PDF字典。 - 访问PDF字典:使用
GetPdfObject()可以直接获取该字体在PDF文件中的字典数据。 - 提取名称:字体名称存储在
BaseFont子字典的Name键下——这正是Acrobat Reader用来显示字体名称的数据源。 - 统一处理逻辑:保留了针对标准字体的原有逻辑,所以这个类现在可以正确捕获所有类型的字体名称。
这种方法能够返回与Acrobat Reader中显示的Type 3字体名称一致的结果,同时也能无缝处理其他类型的字体。
内容的提问来源于stack exchange,提问作者Jules
相关产品推荐
相关产品推荐

