为何出现“无法读取GoogleSans-Regular字体的ToUnicode CMap”错误?
问题:PDFBox解析第三方PDF时ToUnicode CMap报错,寻求通用处理方案
不确定问题出在PDF文件、PDFBox组件还是自身操作上,推测大概率是文件问题。当前遇到报错:Could not read ToUnicode CMap in font GoogleSans-Regular,伴随异常栈信息:
java.io.IOException: java.lang.IllegalArgumentException: The start and the end values must not have different lengths. at org.apache.fontbox.cmap.CMapParser.parseBegincodespacerange(CMapParser.java:289) at org.apache.fontbox.cmap.CMapParser.parse(CMapParser.java:147) at org.apache.pdfbox.pdmodel.font.CMapManager.parseCMap(CMapManager.java:73) at org.apache.pdfbox.pdmodel.font.PDFont.readCMap(PDFont.java:218) at org.apache.pdfbox.pdmodel.font.PDFont.loadUnicodeCmap(PDFont.java:147) at org.apache.pdfbox.pdmodel.font.PDFont.<init>(PDFont.java:115) at org.apache.pdfbox.pdmodel.font.PDType0Font.<init>(PDType0Font.java:182) at org.apache.pdfbox.pdmodel.font.PDFontFactory.createFont(PDFontFactory.java:97) at org.apache.pdfbox.pdmodel.PDResources.getFont(PDResources.java:171) at org.apache.pdfbox.contentstream.operator.text.SetFontAndSize.process(SetFontAndSize.java:66) at org.apache.pdfbox.contentstream.PDFStreamEngine.processOperator(PDFStreamEngine.java:966) at org.apache.pdfbox.contentstream.PDFStreamEngine.processStreamOperators(PDFStreamEngine.java:541)
该PDF的ToUnicode内容如下:
/CIDInit /ProcSet findresource begin 12 dict begin begincmap /CIDSystemInfo << /Registry (Adobe) /Ordering (Identity) /Supplement 0 >> def /CMapName /Adobe-Identity-H def CMapType 2 def 1 begincodespacerange <0000> <FFFFF> endcodespacerange 0 beginbfchar endbfchar 1 beginbfrange <0003> <0037> [<0020> <0041> <0042> <0043> <0044> <0045> <0046> <0047> <0048> <0049> <004A> <004B> <004C> <004D> <004E> <004F> <0050> <0051> <0052> <0053> <0054> <0055> <0056> <0057> <0058> <0059> <005A> <0061> <0062> <0063> <0064> <0065> <0066> <0067> <0068> <0069> <006A> <006B> <006C> <006D> <006E> <006F> <0070> <0071> <0072> <0073> <0074> <0075> <0076> <0077> <0078> <0079> <007A>] endbfrange endcmap CMapName currentdict /CMap defineresource pop end end
需处理大量第三方PDF,希望找到通用处理方案而非修复单个文件,询问是否可让PDFBox默认采用Unicode解析。
解决方案
报错原因分析
从提供的ToUnicode CMap内容看,begincodespacerange中的<0000>是4位十六进制(2字节),<FFFFF>是5位(3字节),违反了PDF规范中代码空间范围的起始和结束值必须长度一致的要求,这是PDF文件本身的非法格式问题。
通用处理方案
方案1:自定义CMapParser做容错处理
继承PDFBox的CMapParser类,重写parseBegincodespacerange方法,对长度不一致的代码范围进行自动修正,比如截断较长值到较短长度,或者填充较短值到较长长度。示例代码:
public class TolerantCMapParser extends CMapParser { public TolerantCMapParser(InputStream input) throws IOException { super(input); } @Override protected void parseBegincodespacerange() throws IOException { try { super.parseBegincodespacerange(); } catch (IllegalArgumentException e) { // 容错处理:取起始和结束值的最小长度截断 List<CodeSpaceRange> ranges = new ArrayList<>(); int count = readInt(); for (int i = 0; i < count; i++) { String startStr = readHexString(); String endStr = readHexString(); int minLength = Math.min(startStr.length(), endStr.length()); startStr = startStr.substring(0, minLength); endStr = endStr.substring(0, minLength); byte[] start = Hex.decodeHex(startStr.toCharArray()); byte[] end = Hex.decodeHex(endStr.toCharArray()); ranges.add(new CodeSpaceRange(start, end)); } getCMap().setCodeSpaceRanges(ranges); } } }
之后需要替换PDFBox默认的CMap解析逻辑,可通过自定义CMapManager或在加载字体时指定使用该容错解析器。
方案2:禁用ToUnicode CMap,强制使用字体编码映射
通过自定义文本剥离器(PDFTextStripper),跳过有问题的ToUnicode CMap,转而使用字体内置编码进行解析。示例代码:
public class TolerantTextStripper extends PDFTextStripper { @Override protected PDFont getFont(COSName fontName) throws IOException { try { return super.getFont(fontName); } catch (IOException e) { // 加载失败时,使用默认字体替代并强制Unicode映射 PDResources resources = getCurrentPage().getResources(); COSDictionary fontDict = resources.getFontDictionary(fontName); // 可根据实际场景选择合适的替代字体 return PDType1Font.HELVETICA; } } }
方案3:升级PDFBox到最新稳定版
旧版本PDFBox对非法CMap的容错性较差,升级到最新稳定版(如2.0.x或3.0.x分支),新版本可能已内置对这类非法格式的自动处理逻辑,无需额外编码即可解决问题。
内容的提问来源于stack exchange,提问作者curiousity
相关产品推荐
相关产品推荐

