PDFBox填充含土耳其字符表单报错:U+015E不在Helvetica字体中
问题描述
我有一个带可编辑字段的简易PDF文件,需要填充信息。使用英文字符时一切正常,但部分字段包含土耳其字符(如Ş)。我尝试过按StackOverflow上的方法添加不同字体(如Arial),也试过PDFBox 2.x版本,均未解决问题。示例代码如下:
public static void main(String[] args) throws IOException { try (PDDocument doc = PDDocument.load(Main.class.getResourceAsStream("/pdf/template.pdf"))) { // Access the AcroForm PDAcroForm acroForm = doc.getDocumentCatalog().getAcroForm(); PDType0Font load = PDType0Font.load(doc, new File("C:/windows/fonts/arial.ttf")); PDResources resources = acroForm.getDefaultResources(); if (resources == null) { resources = new PDResources(); acroForm.setDefaultResources(resources); } String fontName = resources.add(load).getName(); String defaultAppearanceString = "/" + fontName + " 12 Tf 0 g"; acroForm.setDefaultResources(resources); PDTextField myField = (PDTextField) acroForm.getField("field"); myField.setDefaultAppearance(defaultAppearanceString); myField.getWidgets().get(0).setAppearance(null); myField.setValue("ŞŞ"); // Text with the Ş character // Save the updated document doc.save("target/SimpleFormWithCorrectFont.pdf"); } }
运行时报错:U+015E ('Scedilla') is not available in the font Helvetica, encoding: WinAnsiEncoding。请问我的操作存在什么错误?
问题原因与解决方法
错误核心
你只给表单全局设置了默认资源和外观字符串,但没处理字段控件(Widget)的独立资源配置,也没指定支持Unicode的编码,导致PDFBox仍用默认的Helvetica字体(不支持土耳其字符)和WinAnsiEncoding(不覆盖扩展Unicode字符)来渲染内容。
关键修正步骤
- 给字段Widget单独添加字体资源:部分PDF模板的字段控件会自带独立资源集合,优先级高于表单全局设置,必须确保Widget能访问到Arial字体。
- 显式设置字段编码为UTF-16:WinAnsiEncoding仅支持基本ASCII和西欧字符,UTF-16可覆盖所有Unicode字符,包括土耳其特殊字符。
修正后的代码
public static void main(String[] args) throws IOException { try (PDDocument doc = PDDocument.load(Main.class.getResourceAsStream("/pdf/template.pdf"))) { PDAcroForm acroForm = doc.getDocumentCatalog().getAcroForm(); PDType0Font arialFont = PDType0Font.load(doc, new File("C:/windows/fonts/arial.ttf")); // 配置表单全局默认资源 PDResources formResources = acroForm.getDefaultResources(); if (formResources == null) { formResources = new PDResources(); acroForm.setDefaultResources(formResources); } String fontName = formResources.add(arialFont).getName(); String defaultAppearance = "/" + fontName + " 12 Tf 0 g"; acroForm.setDefaultAppearance(defaultAppearance); PDTextField myField = (PDTextField) acroForm.getField("field"); // 配置字段Widget的独立资源(关键) PDResources widgetResources = myField.getWidgets().get(0).getResources(); if (widgetResources == null) { widgetResources = new PDResources(); myField.getWidgets().get(0).setResources(widgetResources); } widgetResources.add(arialFont); // 覆盖字段的默认外观配置 myField.setDefaultAppearance(defaultAppearance); // 强制设置UTF-16编码支持特殊字符 myField.setCOSString(COSName.ENCODING, COSName.UTF_16); // 清除旧外观缓存 myField.getWidgets().get(0).setAppearance(null); // 设置带土耳其字符的内容 myField.setValue("ŞŞ"); doc.save("target/SimpleFormWithCorrectFont.pdf"); } }
补充说明
- 若你的PDF模板存在多个带特殊字符的字段,可循环处理所有字段的Widget资源和编码,避免重复代码。
- 确保加载的Arial字体文件路径正确,Windows系统下也可使用
Font.createFont(Font.TRUETYPE_FONT, new File("C:/windows/fonts/arial.ttf"))来验证字体是否能正常读取。
内容的提问来源于stack exchange,提问作者Omnis
相关产品推荐
相关产品推荐

