Java逐行提取二进制数据时字符丢失问题及适配实现求助
问题根源
使用BufferedReader逐行读取字符串时,会将字节流通过字符编码转换为字符串。如果二进制数据中的字节(如0x81)无法被当前编码(默认通常是UTF-8)识别为有效字符,就会被替换为0x3F(替换字符),导致数据损坏。
解决方案
由于二进制数据不是文本,不能直接用字符流处理。有两种可靠的方式实现按字节逐行处理并保留原始数据:
方案1:利用ISO-8859-1编码的1:1字节映射(简单高效)
ISO-8859-1编码会将每一个字节(0x00-0xFF)映射为唯一的字符,不会丢失任何数据。我们可以用它来读取行,再转换回字节数组:
import java.io.*; import java.nio.charset.StandardCharsets; public class BinaryLineProcessor { public static void processBinaryFile(String inputPath, String outputPath) throws IOException { try (BufferedReader reader = new BufferedReader( new InputStreamReader(new FileInputStream(inputPath), StandardCharsets.ISO_8859_1)); FileOutputStream outputStream = new FileOutputStream(outputPath)) { String line; while ((line = reader.readLine()) != null) { // 将字符串转换回原始字节数组(不含行分隔符) byte[] lineBytes = line.getBytes(StandardCharsets.ISO_8859_1); // 在这里处理行字节:提取TID和二进制数据 byte[] processedData = processLine(lineBytes); // 写入处理后的数据到输出文件 outputStream.write(processedData); // 若需保留行结构,手动补充原文件的行分隔符(如CRLF) // outputStream.write("\r\n".getBytes(StandardCharsets.ISO_8859_1)); } } } private static byte[] processLine(byte[] lineBytes) { // 示例:提取前4字节作为TID,剩余为二进制数据 byte[] tid = new byte[4]; System.arraycopy(lineBytes, 0, tid, 0, 4); byte[] binaryData = new byte[lineBytes.length - 4]; System.arraycopy(lineBytes, 4, binaryData, 0, binaryData.length); // 构造包含TID和二进制数据的输出字节数组 byte[] result = new byte[tid.length + binaryData.length]; System.arraycopy(tid, 0, result, 0, tid.length); System.arraycopy(binaryData, 0, result, tid.length, binaryData.length); return result; } }
方案2:手动按字节读取行(完全控制)
如果需要处理非标准行分隔符,可直接用字节流逐字节读取,手动拼接行:
import java.io.*; public class ManualBinaryLineProcessor { private static final byte LF = 0x0A; private static final byte CR = 0x0D; public static void processBinaryFile(String inputPath, String outputPath) throws IOException { try (BufferedInputStream inputStream = new BufferedInputStream(new FileInputStream(inputPath)); FileOutputStream outputStream = new FileOutputStream(outputPath)) { ByteArrayOutputStream lineBuffer = new ByteArrayOutputStream(); int currentByte; boolean prevWasCR = false; while ((currentByte = inputStream.read()) != -1) { if (currentByte == LF) { // 处理当前行 byte[] lineBytes = lineBuffer.toByteArray(); byte[] processedData = processLine(lineBytes); outputStream.write(processedData); lineBuffer.reset(); prevWasCR = false; } else if (currentByte == CR) { prevWasCR = true; } else { if (prevWasCR) { lineBuffer.write(CR); prevWasCR = false; } lineBuffer.write(currentByte); } } // 处理文件末尾未换行的最后一行 if (lineBuffer.size() > 0) { byte[] lineBytes = lineBuffer.toByteArray(); byte[] processedData = processLine(lineBytes); outputStream.write(processedData); } } } private static byte[] processLine(byte[] lineBytes) { // 同方案1的处理逻辑,按需调整 byte[] tid = new byte[4]; System.arraycopy(lineBytes, 0, tid, 0, 4); byte[] binaryData = new byte[lineBytes.length - 4]; System.arraycopy(lineBytes, 4, binaryData, 0, binaryData.length); byte[] result = new byte[tid.length + binaryData.length]; System.arraycopy(tid, 0, result, 0, tid.length); System.arraycopy(binaryData, 0, result, tid.length, binaryData.length); return result; } }
关键注意事项
- 输出必须用
FileOutputStream直接写入字节,禁止使用字符流(如PrintWriter),避免二次编码损坏数据。 readLine()会自动去掉行分隔符(CR/LF),若需保留原文件的行结构,需手动补充对应的分隔符。- 方案1的
ISO-8859-1技巧仅适用于字节与字符1:1映射的场景,不能用于UTF-8等可变长度编码。
内容的提问来源于stack exchange,提问作者Edisson Gabriel López
相关产品推荐
相关产品推荐

