Java中如何读取包含byte array与List array的未知内容文件?
当未知文件结构时读取混合byte数组与字符串列表的方法
如果完全不清楚文件的内容结构,直接自动区分并读取其中的byte数组和字符串列表没有通用的完美解法——因为二进制数据本身不具备自描述性,无法仅凭字节流判断哪部分是原始byte数组、哪部分是字符串。但可以根据写入时可能用到的规则(比如你提到的DataOutputStream),尝试以下几种可行方案:
1. 基于DataOutputStream特征的试探性读取
如果文件是通过DataOutputStream写入的,内容会遵循它的格式规则:
writeInt会写入4字节的带符号整数;writeUTF会先写入2字节的长度(表示后续modified UTF-8编码的字节数),再写入字符串的编码字节。
可以基于这些规则,尝试常见的结构组合进行试探:
import java.io.*; import java.util.ArrayList; import java.util.Arrays; import java.util.List; public class UnknownFormatReader { public static void main(String[] args) { try (FileInputStream fis = new FileInputStream("h.txt"); DataInputStream dis = new DataInputStream(fis)) { // 尝试第一种结构:先byte数组,再字符串列表 if (tryReadByteFirst(dis)) { return; } // 重置流,尝试第二种结构:先字符串列表,再byte数组 fis.getChannel().position(0); if (tryReadStringFirst(dis)) { return; } System.err.println("无法识别文件的结构"); } catch (IOException e) { e.printStackTrace(); } } private static boolean tryReadByteFirst(DataInputStream dis) throws IOException { try { int byteLen = dis.readInt(); byte[] bytes = new byte[byteLen]; dis.readFully(bytes); int strCount = dis.readInt(); List<String> strings = new ArrayList<>(); for (int i = 0; i < strCount; i++) { strings.add(dis.readUTF()); } // 验证是否读到文件末尾 if (dis.available() == 0) { System.out.println("读取成功(先byte数组):"); System.out.println("Byte数组:" + Arrays.toString(bytes)); System.out.println("字符串列表:" + strings); return true; } } catch (EOFException | UTFDataFormatException e) { // 结构不匹配,忽略异常 } return false; } private static boolean tryReadStringFirst(DataInputStream dis) throws IOException { try { int strCount = dis.readInt(); List<String> strings = new ArrayList<>(); for (int i = 0; i < strCount; i++) { strings.add(dis.readUTF()); } int byteLen = dis.readInt(); byte[] bytes = new byte[byteLen]; dis.readFully(bytes); if (dis.available() == 0) { System.out.println("读取成功(先字符串列表):"); System.out.println("字符串列表:" + strings); System.out.println("Byte数组:" + Arrays.toString(bytes)); return true; } } catch (EOFException | UTFDataFormatException e) { // 结构不匹配,忽略异常 } return false; } }
这种方法的核心是:如果读取过程中没有抛出异常(比如字节不足、UTF格式错误),且刚好读到文件末尾,说明结构匹配。
2. 分析二进制特征手动识别
如果试探法失败,可以用二进制编辑器打开文件,分析字节特征:
- 寻找可能的长度标记:4字节的整数(对应
writeInt写入的数组长度),观察后续字节的长度是否和该整数匹配; - 寻找
writeUTF的特征:2字节的短整数(字符串编码长度),后续字节尝试解码为UTF-8字符串,如果能正常解码,说明这部分是字符串; - 原始byte数组没有固定特征,只能通过排除法判断——排除掉符合字符串格式的部分后,剩余的就是原始byte数组。
3. 从根源解决:写入时添加自描述元数据
如果是你自己控制文件的写入流程,最好在文件开头添加结构描述信息,让读取方不需要提前了解结构就能解析。比如:
写入示例(添加结构标记)
import java.io.*; import java.util.ArrayList; import java.util.List; public class StructuredWriter { public static void main(String[] args) throws IOException { byte[] bytes = {1, 2, 3, 4, 5}; List<String> strings = new ArrayList<>(); strings.add("Hello"); strings.add("World"); strings.add("!"); try (FileOutputStream fos = new FileOutputStream("h.txt"); DataOutputStream dos = new DataOutputStream(fos)) { // 写入结构标记:1代表[byte数组][字符串列表],2代表[字符串列表][byte数组] dos.writeByte(1); // 写入byte数组 dos.writeInt(bytes.length); dos.write(bytes); // 写入字符串列表 dos.writeInt(strings.size()); for (String s : strings) { dos.writeUTF(s); } } } }
读取示例(根据标记解析)
import java.io.*; import java.util.ArrayList; import java.util.Arrays; import java.util.List; public class StructuredReader { public static void main(String[] args) throws IOException { try (FileInputStream fis = new FileInputStream("h.txt"); DataInputStream dis = new DataInputStream(fis)) { byte structureFlag = dis.readByte(); if (structureFlag == 1) { // 先读byte数组 int byteLen = dis.readInt(); byte[] bytes = new byte[byteLen]; dis.readFully(bytes); // 再读字符串列表 int strCount = dis.readInt(); List<String> strings = new ArrayList<>(); for (int i = 0; i < strCount; i++) { strings.add(dis.readUTF()); } System.out.println("Byte数组:" + Arrays.toString(bytes)); System.out.println("字符串列表:" + strings); } else if (structureFlag == 2) { // 先读字符串列表 int strCount = dis.readInt(); List<String> strings = new ArrayList<>(); for (int i = 0; i < strCount; i++) { strings.add(dis.readUTF()); } // 再读byte数组 int byteLen = dis.readInt(); byte[] bytes = new byte[byteLen]; dis.readFully(bytes); System.out.println("字符串列表:" + strings); System.out.println("Byte数组:" + Arrays.toString(bytes)); } else { System.err.println("未知的结构标记"); } } } }
这种方法是最可靠的,完全不需要提前了解文件结构,所有信息都包含在文件本身中。
内容的提问来源于stack exchange,提问作者Hamed
相关产品推荐
相关产品推荐

