You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

为何出现“无法读取GoogleSans-Regular字体的ToUnicode CMap”错误?

问题:PDFBox解析第三方PDF时ToUnicode CMap报错,寻求通用处理方案

不确定问题出在PDF文件、PDFBox组件还是自身操作上,推测大概率是文件问题。当前遇到报错:Could not read ToUnicode CMap in font GoogleSans-Regular,伴随异常栈信息:

java.io.IOException: java.lang.IllegalArgumentException: The start and the end values must not have different lengths.
at org.apache.fontbox.cmap.CMapParser.parseBegincodespacerange(CMapParser.java:289)
at org.apache.fontbox.cmap.CMapParser.parse(CMapParser.java:147)
at org.apache.pdfbox.pdmodel.font.CMapManager.parseCMap(CMapManager.java:73)
at org.apache.pdfbox.pdmodel.font.PDFont.readCMap(PDFont.java:218)
at org.apache.pdfbox.pdmodel.font.PDFont.loadUnicodeCmap(PDFont.java:147)
at org.apache.pdfbox.pdmodel.font.PDFont.<init>(PDFont.java:115)
at org.apache.pdfbox.pdmodel.font.PDType0Font.<init>(PDType0Font.java:182)
at org.apache.pdfbox.pdmodel.font.PDFontFactory.createFont(PDFontFactory.java:97)
at org.apache.pdfbox.pdmodel.PDResources.getFont(PDResources.java:171)
at org.apache.pdfbox.contentstream.operator.text.SetFontAndSize.process(SetFontAndSize.java:66)
at org.apache.pdfbox.contentstream.PDFStreamEngine.processOperator(PDFStreamEngine.java:966)
at org.apache.pdfbox.contentstream.PDFStreamEngine.processStreamOperators(PDFStreamEngine.java:541)

该PDF的ToUnicode内容如下:

/CIDInit /ProcSet findresource begin
12 dict begin
begincmap
/CIDSystemInfo
<< /Registry (Adobe)
/Ordering (Identity)
/Supplement 0
>> def
/CMapName /Adobe-Identity-H def
CMapType 2 def
1 begincodespacerange
<0000> <FFFFF>
endcodespacerange
0 beginbfchar
endbfchar
1 beginbfrange
<0003> <0037> [<0020> <0041> <0042> <0043> <0044> <0045> <0046> <0047> <0048> <0049> <004A> <004B> <004C> <004D> <004E> <004F> <0050> <0051> <0052> <0053> <0054> <0055> <0056> <0057> <0058> <0059> <005A> <0061> <0062> <0063> <0064> <0065> <0066> <0067> <0068> <0069> <006A> <006B> <006C> <006D> <006E> <006F> <0070> <0071> <0072> <0073> <0074> <0075> <0076> <0077> <0078> <0079> <007A>]
endbfrange
endcmap
CMapName currentdict /CMap defineresource pop
end
end

需处理大量第三方PDF,希望找到通用处理方案而非修复单个文件,询问是否可让PDFBox默认采用Unicode解析。


解决方案

报错原因分析

从提供的ToUnicode CMap内容看,begincodespacerange中的<0000>是4位十六进制(2字节),<FFFFF>是5位(3字节),违反了PDF规范中代码空间范围的起始和结束值必须长度一致的要求,这是PDF文件本身的非法格式问题。

通用处理方案

方案1:自定义CMapParser做容错处理

继承PDFBox的CMapParser类,重写parseBegincodespacerange方法,对长度不一致的代码范围进行自动修正,比如截断较长值到较短长度,或者填充较短值到较长长度。示例代码:

public class TolerantCMapParser extends CMapParser {
    public TolerantCMapParser(InputStream input) throws IOException {
        super(input);
    }

    @Override
    protected void parseBegincodespacerange() throws IOException {
        try {
            super.parseBegincodespacerange();
        } catch (IllegalArgumentException e) {
            // 容错处理:取起始和结束值的最小长度截断
            List<CodeSpaceRange> ranges = new ArrayList<>();
            int count = readInt();
            for (int i = 0; i < count; i++) {
                String startStr = readHexString();
                String endStr = readHexString();
                int minLength = Math.min(startStr.length(), endStr.length());
                startStr = startStr.substring(0, minLength);
                endStr = endStr.substring(0, minLength);
                byte[] start = Hex.decodeHex(startStr.toCharArray());
                byte[] end = Hex.decodeHex(endStr.toCharArray());
                ranges.add(new CodeSpaceRange(start, end));
            }
            getCMap().setCodeSpaceRanges(ranges);
        }
    }
}

之后需要替换PDFBox默认的CMap解析逻辑,可通过自定义CMapManager或在加载字体时指定使用该容错解析器。

方案2:禁用ToUnicode CMap,强制使用字体编码映射

通过自定义文本剥离器(PDFTextStripper),跳过有问题的ToUnicode CMap,转而使用字体内置编码进行解析。示例代码:

public class TolerantTextStripper extends PDFTextStripper {
    @Override
    protected PDFont getFont(COSName fontName) throws IOException {
        try {
            return super.getFont(fontName);
        } catch (IOException e) {
            // 加载失败时,使用默认字体替代并强制Unicode映射
            PDResources resources = getCurrentPage().getResources();
            COSDictionary fontDict = resources.getFontDictionary(fontName);
            // 可根据实际场景选择合适的替代字体
            return PDType1Font.HELVETICA;
        }
    }
}

方案3:升级PDFBox到最新稳定版

旧版本PDFBox对非法CMap的容错性较差,升级到最新稳定版(如2.0.x或3.0.x分支),新版本可能已内置对这类非法格式的自动处理逻辑,无需额外编码即可解决问题。


内容的提问来源于stack exchange,提问作者curiousity

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.13 00:37:46