如何将原始PDF文件字符串转Base64?解码后PDF无法生成求助
PDF字符串转Base64后无法正常解码的问题解决
问题背景
尝试将从其他服务获取的原始PDF字符串转换为Base64,生成的Base64格式看似正确,但解码后无法生成可正常打开的PDF文件。已尝试UTF-8编码等多种转换方法,因无法获取实际PDF文件,只能处理该原始字符串。
原始代码
public String get() throws CustomException, IOException { String testPDf = "%PDF-1.4\n" + "%âãÏÓ\n" + "1 0 obj\n" + "<< /Type /Catalog\n" + " /Pages 2 0 R\n" + ">>\n" + "endobj\n" + "2 0 obj\n" + "<< /Type /Pages\n" + " /Kids [3 0 R]\n" + " /Count 1\n" + ">>\n" + "endobj\n" + "3 0 obj\n" + "<< /Type /Page\n" + " /Parent 2 0 R\n" + " /Resources << /Font << /F1 << /Type /Font\n" + " /Subtype /Type1\n" + " /BaseFont /Helvetica\n" + " >> >>\n" + " /MediaBox [0 0 612 792]\n" + " /Contents 4 0 R\n" + ">>\n" + "endobj\n" + "4 0 obj\n" + "<< /Length 55 >>\n" + "stream\n" + "BT\n" + "/F1 18 Tf\n" + "100 100 Td\n" + "(Hello, World!) Tj\n" + "ET\n" + "endstream\n" + "endobj\n" + "xref\n" + "0 5\n" + "0000000000 65535 f\n" + "0000000018 00000 n\n" + "0000000077 00000 n\n" + "0000000175 00000 n\n" + "0000000451 00000 n\n" + "trailer\n" + "<< /Size 5\n" + " /Root 1 0 R\n" + ">>\n" + "startxref\n" + "561\n" + "%%EOF"; testPDf = testPDf.replaceAll("\n", ""); testPDf = toBase64(testPDf); return testPDf; } public static String toBase64(String str) throws UnsupportedEncodingException { byte[] bytes = str.getBytes("UTF-8"); String encoded = Base64.getEncoder().encodeToString(bytes); return encoded; }
错误原因
- 破坏PDF结构:代码中通过
replaceAll("\n", "")删除了所有换行符,但PDF是结构化的二进制文件,其xref表、startxref的偏移量均基于原始字节流位置计算,去掉换行后会导致解析器无法识别文件结构。 - 错误的编码方式:PDF并非纯UTF-8文本,包含二进制字节(比如文件头的
%âãÏÓ是PDF的二进制标记),用UTF-8编码转换字符串会改变原始字节,导致数据损坏。
解决方法
- 保留原始换行符:不要去除PDF字符串中的换行,维持文件的原始结构。
- 使用单字节编码还原原始字节:采用
ISO-8859-1编码(单字节编码,可无损保留所有原始字节)将字符串转换为字节数组,再进行Base64编码。
修正后的代码
public String get() throws CustomException, IOException { String testPDf = "%PDF-1.4\n" + "%âãÏÓ\n" + "1 0 obj\n" + "<< /Type /Catalog\n" + " /Pages 2 0 R\n" + ">>\n" + "endobj\n" + "2 0 obj\n" + "<< /Type /Pages\n" + " /Kids [3 0 R]\n" + " /Count 1\n" + ">>\n" + "endobj\n" + "3 0 obj\n" + "<< /Type /Page\n" + " /Parent 2 0 R\n" + " /Resources << /Font << /F1 << /Type /Font\n" + " /Subtype /Type1\n" + " /BaseFont /Helvetica\n" + " >> >>\n" + " /MediaBox [0 0 612 792]\n" + " /Contents 4 0 R\n" + ">>\n" + "endobj\n" + "4 0 obj\n" + "<< /Length 55 >>\n" + "stream\n" + "BT\n" + "/F1 18 Tf\n" + "100 100 Td\n" + "(Hello, World!) Tj\n" + "ET\n" + "endstream\n" + "endobj\n" + "xref\n" + "0 5\n" + "0000000000 65535 f\n" + "0000000018 00000 n\n" + "0000000077 00000 n\n" + "0000000175 00000 n\n" + "0000000451 00000 n\n" + "trailer\n" + "<< /Size 5\n" + " /Root 1 0 R\n" + ">>\n" + "startxref\n" + "561\n" + "%%EOF"; // 移除错误的换行符删除操作 testPDf = toBase64(testPDf); return testPDf; } public static String toBase64(String str) throws UnsupportedEncodingException { // 使用ISO-8859-1编码还原原始字节 byte[] bytes = str.getBytes("ISO-8859-1"); String encoded = Base64.getEncoder().encodeToString(bytes); return encoded; }
内容的提问来源于stack exchange,提问作者BobbySin
相关产品推荐
相关产品推荐

