You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用iText7获取PDF中Type 3字体的实际名称?

获取iText7中Type 3字体的名称

Great question! I’ve dealt with this exact quirk of Type 3 fonts in iText7 before—they don’t expose their names through the standard GetFontNames() method because they’re embedded as custom, user-defined fonts rather than pre-defined font programs. Acrobat Reader pulls the name directly from the underlying PDF dictionary, and we can do the same in C#.

问题原因

Type 3字体在iText7中由Type3Font实例表示,它们的FontProgram(即Type3FontProgram)不会填充GetFontNames()所需的字体名称数据。相反,字体名称存储在PDF的BaseFont字典条目中的Name键下。

修改后的解决方案代码

下面是更新后的FontReader类代码,它可以同时捕获标准字体和Type 3字体的名称:

private class FontReader : IEventListener {
    public ICollection<string> Fonts { get; }
    public FontReader() {
        Fonts = new List<string>();
    }
    public void EventOccurred(IEventData data, EventType type) {
        if (!(data is TextRenderInfo textRenderInfo)) return;
        
        var font = textRenderInfo.GetFont();
        string fontName = null;

        // 处理标准字体类型(Type 1、TrueType等)
        var fontProgram = font.GetFontProgram();
        if (!(fontProgram is Type3FontProgram)) {
            fontName = fontProgram.GetFontNames().GetFontName();
        }
        // 通过直接读取PDF字典处理Type 3字体
        else if (font is Type3Font type3Font) {
            var fontDict = type3Font.GetPdfObject() as PdfDictionary;
            if (fontDict != null) {
                var baseFont = fontDict.GetAsDictionary(PdfName.BaseFont);
                if (baseFont != null) {
                    var nameEntry = baseFont.GetAsString(PdfName.Name);
                    fontName = nameEntry?.GetValue();
                }
            }
        }

        // 如果名称有效且未在集合中,则添加
        if (!string.IsNullOrEmpty(fontName) && !Fonts.Contains(fontName)) {
            Fonts.Add(fontName);
        }
    }
    public ICollection<EventType> GetSupportedEvents() {
        return new HashSet<EventType> {EventType.RENDER_TEXT};
    }
}

关键步骤说明

  • 判断Type 3字体:首先验证字体是否为Type3Font实例,这意味着我们需要访问底层的PDF字典。
  • 访问PDF字典:使用GetPdfObject()可以直接获取该字体在PDF文件中的字典数据。
  • 提取名称:字体名称存储在BaseFont子字典的Name键下——这正是Acrobat Reader用来显示字体名称的数据源。
  • 统一处理逻辑:保留了针对标准字体的原有逻辑,所以这个类现在可以正确捕获所有类型的字体名称。

这种方法能够返回与Acrobat Reader中显示的Type 3字体名称一致的结果,同时也能无缝处理其他类型的字体。

内容的提问来源于stack exchange,提问作者Jules

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.07 07:59:07