从CSV读取Blob转存PDF无法打开,求解决方案
问题:Blob转PDF文件无法打开,提示无法将application/octet-stream转为application/pdf
从CSV读取数据存入名为record的Map对象,尝试将其中的Blob数据转换保存为文档,图片能正常打开,但PDF保存后无法打开报错。在线工具转换时提示无法将application/octet-stream转换为application/pdf。
相关代码
private static void convertBlobToBase64(Map<String, Object> record) { boolean blob = record.keySet().stream().anyMatch(key -> key.contains("Blob")); if (blob) { for (Map.Entry<String, Object> entry : record.entrySet()) { if (entry.getKey().contains("Blob")) { try { String hexData = (String) entry.getValue(); hexData = hexData.substring(1).replaceAll("[^0-9A-Fa-f]", ""); if (hexData.length() % 2 != 0) { throw new IllegalArgumentException("Invalid hex string"); } byte[] hexBinary = DatatypeConverter.parseHexBinary(hexData); String base64String = Base64.getEncoder().encodeToString(hexBinary); blobToDoc(base64String, (String) record.get("Document_FileName")); } catch (IOException | SQLException e) { throw new RuntimeException("Error converting blob: "+ e.getMessage()); } return; } } } } private static void blobToDoc(String blobData, String fileName) throws SQLException, IOException { byte[] blobValue = Base64.getDecoder().decode(blobData); Blob blob = new SerialBlob(blobValue); String path = "src/main/resources/" + fileName; Path outputFile = Paths.get(path); try (OutputStream outputStream = Files.newOutputStream(outputFile)) { byte[] buffer = new byte[4096]; int bytesRead = -1; InputStream inputStream = blob.getBinaryStream(); while ((bytesRead = inputStream.read(buffer)) != -1) { outputStream.write(buffer, 0, bytesRead); } inputStream.close(); outputStream.flush(); } }
错误情况
保存后的PDF文件无法打开,在线工具转换时提示无法将application/octet-stream类型数据转换为application/pdf。
解决方法
1. 确认Hex数据解析逻辑的正确性
- 检查
hexData.substring(1)操作:如果原始Hex字符串没有多余前缀(比如0x),这一步会丢失第一个字节的数据,直接导致PDF文件损坏。请根据实际数据格式决定是否保留这行代码。 - 验证
replaceAll("[^0-9A-Fa-f]", "")的过滤结果:确保没有误删有效十六进制字符,导致解析后的字节数组和原始Blob数据不一致。
2. 简化转换流程,减少冗余步骤
当前代码存在多次不必要的转换(Hex→Byte→Base64→Byte→Blob→文件),中间环节容易引入错误。直接将解析后的字节数组写入文件即可:
private static void convertBlobToFile(Map<String, Object> record) { boolean hasBlob = record.keySet().stream().anyMatch(key -> key.contains("Blob")); if (!hasBlob) return; for (Map.Entry<String, Object> entry : record.entrySet()) { if (entry.getKey().contains("Blob")) { try { String hexData = (String) entry.getValue(); // 仅当原始Hex有0x前缀时才执行此操作,否则注释 // hexData = hexData.substring(1); hexData = hexData.replaceAll("[^0-9A-Fa-f]", ""); if (hexData.length() % 2 != 0) { throw new IllegalArgumentException("Invalid hex string"); } byte[] fileBytes = DatatypeConverter.parseHexBinary(hexData); writeBytesToFile(fileBytes, (String) record.get("Document_FileName")); } catch (IOException e) { throw new RuntimeException("Error converting blob: " + e.getMessage()); } return; } } } private static void writeBytesToFile(byte[] fileBytes, String fileName) throws IOException { String path = "src/main/resources/" + fileName; Path outputFile = Paths.get(path); Files.write(outputFile, fileBytes); }
3. 验证原始数据是否为有效PDF
如果修改后仍无法打开,需确认CSV中的Blob数据本身是否有效:
- 取解析后的字节数组前5个字节,对比PDF文件的魔数
%PDF-(对应十六进制:25 50 44 46 2D),如果不匹配,说明原始数据不是PDF,或在CSV存储过程中已损坏。 - 检查CSV文件中Blob字段的格式,是否存在换行、转义字符等导致数据截断或篡改。
关于application/octet-stream的说明
application/octet-stream只是通用二进制文件的MIME类型,并非数据本身不能是PDF。只要原始二进制数据是有效的PDF,直接保存为.pdf后缀的文件即可正常打开,不存在“转换”的说法——问题根源在于你的数据解析或转换流程中丢失/篡改了原始PDF字节。
内容的提问来源于stack exchange,提问作者Ikechukwu Paul Anene
相关产品推荐
相关产品推荐

