You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Java逐行提取二进制数据时字符丢失问题及适配实现求助

问题根源

使用BufferedReader逐行读取字符串时,会将字节流通过字符编码转换为字符串。如果二进制数据中的字节(如0x81)无法被当前编码(默认通常是UTF-8)识别为有效字符,就会被替换为0x3F(替换字符),导致数据损坏。

解决方案

由于二进制数据不是文本,不能直接用字符流处理。有两种可靠的方式实现按字节逐行处理并保留原始数据:

方案1:利用ISO-8859-1编码的1:1字节映射(简单高效)

ISO-8859-1编码会将每一个字节(0x00-0xFF)映射为唯一的字符,不会丢失任何数据。我们可以用它来读取行,再转换回字节数组:

import java.io.*;
import java.nio.charset.StandardCharsets;

public class BinaryLineProcessor {
    public static void processBinaryFile(String inputPath, String outputPath) throws IOException {
        try (BufferedReader reader = new BufferedReader(
                new InputStreamReader(new FileInputStream(inputPath), StandardCharsets.ISO_8859_1));
             FileOutputStream outputStream = new FileOutputStream(outputPath)) {

            String line;
            while ((line = reader.readLine()) != null) {
                // 将字符串转换回原始字节数组(不含行分隔符)
                byte[] lineBytes = line.getBytes(StandardCharsets.ISO_8859_1);
                
                // 在这里处理行字节:提取TID和二进制数据
                byte[] processedData = processLine(lineBytes);
                
                // 写入处理后的数据到输出文件
                outputStream.write(processedData);
                // 若需保留行结构,手动补充原文件的行分隔符(如CRLF)
                // outputStream.write("\r\n".getBytes(StandardCharsets.ISO_8859_1));
            }
        }
    }

    private static byte[] processLine(byte[] lineBytes) {
        // 示例:提取前4字节作为TID,剩余为二进制数据
        byte[] tid = new byte[4];
        System.arraycopy(lineBytes, 0, tid, 0, 4);
        byte[] binaryData = new byte[lineBytes.length - 4];
        System.arraycopy(lineBytes, 4, binaryData, 0, binaryData.length);
        
        // 构造包含TID和二进制数据的输出字节数组
        byte[] result = new byte[tid.length + binaryData.length];
        System.arraycopy(tid, 0, result, 0, tid.length);
        System.arraycopy(binaryData, 0, result, tid.length, binaryData.length);
        
        return result;
    }
}

方案2:手动按字节读取行(完全控制)

如果需要处理非标准行分隔符,可直接用字节流逐字节读取,手动拼接行:

import java.io.*;

public class ManualBinaryLineProcessor {
    private static final byte LF = 0x0A;
    private static final byte CR = 0x0D;

    public static void processBinaryFile(String inputPath, String outputPath) throws IOException {
        try (BufferedInputStream inputStream = new BufferedInputStream(new FileInputStream(inputPath));
             FileOutputStream outputStream = new FileOutputStream(outputPath)) {

            ByteArrayOutputStream lineBuffer = new ByteArrayOutputStream();
            int currentByte;
            boolean prevWasCR = false;

            while ((currentByte = inputStream.read()) != -1) {
                if (currentByte == LF) {
                    // 处理当前行
                    byte[] lineBytes = lineBuffer.toByteArray();
                    byte[] processedData = processLine(lineBytes);
                    outputStream.write(processedData);
                    lineBuffer.reset();
                    prevWasCR = false;
                } else if (currentByte == CR) {
                    prevWasCR = true;
                } else {
                    if (prevWasCR) {
                        lineBuffer.write(CR);
                        prevWasCR = false;
                    }
                    lineBuffer.write(currentByte);
                }
            }

            // 处理文件末尾未换行的最后一行
            if (lineBuffer.size() > 0) {
                byte[] lineBytes = lineBuffer.toByteArray();
                byte[] processedData = processLine(lineBytes);
                outputStream.write(processedData);
            }
        }
    }

    private static byte[] processLine(byte[] lineBytes) {
        // 同方案1的处理逻辑,按需调整
        byte[] tid = new byte[4];
        System.arraycopy(lineBytes, 0, tid, 0, 4);
        byte[] binaryData = new byte[lineBytes.length - 4];
        System.arraycopy(lineBytes, 4, binaryData, 0, binaryData.length);
        
        byte[] result = new byte[tid.length + binaryData.length];
        System.arraycopy(tid, 0, result, 0, tid.length);
        System.arraycopy(binaryData, 0, result, tid.length, binaryData.length);
        
        return result;
    }
}
关键注意事项
  • 输出必须用FileOutputStream直接写入字节,禁止使用字符流(如PrintWriter),避免二次编码损坏数据。
  • readLine()会自动去掉行分隔符(CR/LF),若需保留原文件的行结构,需手动补充对应的分隔符。
  • 方案1的ISO-8859-1技巧仅适用于字节与字符1:1映射的场景,不能用于UTF-8等可变长度编码。

内容的提问来源于stack exchange,提问作者Edisson Gabriel López

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.17 12:15:44