Java复制XLSX文件失败:文件损坏问题排查与解决
解决XLSX文件复制后损坏的问题
问题描述
我编写了一段Java代码测试XLSX文件复制,通过FileWriter和PrintWriter读取源文件Classeur1.xlsx的字节并转为字符串后写入Classeur2.xlsx。源文件可正常打开,但生成的目标文件损坏无法打开。运行输出显示两个文件的字节数组从第15字节开始存在差异,怀疑是编码问题,但不清楚源文件编码(来源多样,通过邮件发送),该如何解决?
测试代码
public static void main(String[] args) throws Exception { FileWriter fileWriter = new FileWriter("Classeur2.xlsx"); PrintWriter printWriter = new PrintWriter(fileWriter); printWriter.print(new String(Files.readAllBytes(Paths.get("Classeur1.xlsx")))); printWriter.close(); byte[] bytes = new byte[150]; System.out.println("++ Classeur1.xlsx ++++++++++++++++++++++"); System.out.println(new String(Files.readAllBytes(Paths.get("Classeur1.xlsx"))).substring(0,150)); System.arraycopy(Files.readAllBytes(Paths.get("Classeur1.xlsx")), 0, bytes, 0, 150); System.out.println(Arrays.toString(bytes)); System.out.println("++ Classeur2.xlsx ++++++++++++++++++++++"); System.out.println(new String(Files.readAllBytes(Paths.get("Classeur2.xlsx"))).substring(0,150)); System.arraycopy(Files.readAllBytes(Paths.get("Classeur2.xlsx")), 0, bytes, 0, 150); System.out.println(Arrays.toString(bytes)); }
运行输出
++ Classeur1.xlsx ++++++++++++++++++++++ PK ! b�h^ � [Content_Types].xml �(� [80, 75, 3, 4, 20, 0, 6, 0, 8, 0, 0, 0, 33, 0, 98, -18, -99, 104, 94, 1, 0, 0, -112, 4, 0, 0, 19, 0, 8, 2, 91, 67, 111, 110, 116, 101, 110, 116, 95, 84, 121, 112, 101, 115, 93, 46, 120, 109, 108, 32, -94, 4, 2, 40, -96, 0, 2, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0] ++ Classeur2.xlsx ++++++++++++++++++++++ PK ! b�h^ � [Content_Types].xml �(� [80, 75, 3, 4, 20, 0, 6, 0, 8, 0, 0, 0, 33, 0, 98, -17, -65, -67, 104, 94, 1, 0, 0, -17, -65, -67, 4, 0, 0, 19, 0, 8, 2, 91, 67, 111, 110, 116, 101, 110, 116, 95, 84, 121, 112, 101, 115, 93, 46, 120, 109, 108, 32, -17, -65, -67, 4, 2, 40, -17, -65, -67, 0, 2, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0]
问题根源
- XLSX是二进制格式:XLSX本质是ZIP压缩包,属于二进制文件,并非文本文件,不能用字符流(
FileWriter/PrintWriter)处理。 - 字节转字符串的编码丢失:将二进制字节转为
String时,JVM会使用默认编码(如UTF-8)解码,遇到无法识别的字节会替换成U+FFFD(即显示的�);写入时这个字符会被编码为UTF-8的三个字节-17,-65,-67,直接篡改了原二进制数据,导致文件损坏。 - 从输出可见的篡改:原文件中的
-18,-99等字节被替换成-17,-65,-67,这就是文件无法打开的直接原因。
解决方法
直接使用字节流复制二进制文件,确保字节原样复制,无需考虑编码:
方式1:使用Files.copy(最简方案)
import java.nio.file.Files; import java.nio.file.Paths; public class XLSXCopy { public static void main(String[] args) throws Exception { Files.copy(Paths.get("Classeur1.xlsx"), Paths.get("Classeur2.xlsx")); } }
方式2:手动字节流读写
import java.io.FileInputStream; import java.io.FileOutputStream; import java.io.IOException; public class XLSXCopy { public static void main(String[] args) throws IOException { try (FileInputStream fis = new FileInputStream("Classeur1.xlsx"); FileOutputStream fos = new FileOutputStream("Classeur2.xlsx")) { byte[] buffer = new byte[4096]; // 4KB缓冲区,平衡性能与内存占用 int bytesRead; while ((bytesRead = fis.read(buffer)) != -1) { fos.write(buffer, 0, bytesRead); } } } }
关键注意事项
- 任何二进制文件(XLSX、PDF、图片、压缩包等)都必须用字节流处理,绝对不能用字符流。
- 二进制文件不存在"文本编码"的概念,直接按字节原样复制即可,无需关心来源文件的编码。
内容的提问来源于stack exchange,提问作者tweetysat
相关产品推荐
相关产品推荐

