You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何移除BufferedInputStream中的{}字符以进行文件编码检测?

如何移除BufferedInputStream中的指定字符后用于编码检测?

我懂你的痛点——字节流可不像字符串那样有现成的replace方法,而且咱们要先处理流再检测编码,还不能提前把字节转成字符(不然就预设了编码,那检测还有啥意义对吧?)。这里有两个实用方案,你可以根据文件大小来选:

方案1:自定义过滤InputStream(适合大文件)

我们可以继承FilterInputStream,重写读取方法,直接跳过{和}对应的字节。这里要说明下:在UTF-8、GBK、ISO-8859-1这类常见编码里,{对应的字节是0x7B,}是0x7D,都是单字节,所以过滤这两个值就能覆盖绝大多数场景。

先写一个自定义的过滤流类:

import java.io.FilterInputStream;
import java.io.IOException;
import java.io.InputStream;

public class CharFilterInputStream extends FilterInputStream {
    // 定义{和}对应的字节值
    private static final byte LEFT_BRACE = 0x7B;
    private static final byte RIGHT_BRACE = 0x7D;

    protected CharFilterInputStream(InputStream in) {
        super(in);
    }

    // 重写单个字节读取方法,跳过目标字符
    @Override
    public int read() throws IOException {
        int b;
        do {
            b = super.read();
        } while (b == LEFT_BRACE || b == RIGHT_BRACE);
        return b;
    }

    // 重写批量读取方法,过滤掉目标字节后返回
    @Override
    public int read(byte[] b, int off, int len) throws IOException {
        int bytesRead = super.read(b, off, len);
        if (bytesRead == -1) {
            return -1;
        }
        int newIndex = off;
        for (int i = off; i < off + bytesRead; i++) {
            if (b[i] != LEFT_BRACE && b[i] != RIGHT_BRACE) {
                b[newIndex++] = b[i];
            }
        }
        return newIndex - off;
    }
}

然后修改你的编码检测方法,把原有的BufferedInputStream包装到这个自定义流里就行:

static String detectCharset(File file) {
    try (FileInputStream fileInputStream = new FileInputStream(file);
         BufferedInputStream bufferedInputStream = new BufferedInputStream(fileInputStream);
         CharFilterInputStream filteredInputStream = new CharFilterInputStream(bufferedInputStream)) {
        
        CharsetDetector charsetDetector = new CharsetDetector();
        charsetDetector.setText(filteredInputStream);
        charsetDetector.enableInputFilter(true);
        CharsetMatch cm = charsetDetector.detect();
        return cm != null ? cm.getName() : null;
    } catch (Exception e) {
        e.printStackTrace();
        return null;
    }
}

方案2:读取到字节数组后过滤(适合小文件)

如果你的文件不大,直接把所有字节读到内存里过滤会更简单,代码量也少:

import java.util.Arrays;

static String detectCharset(File file) {
    try (FileInputStream fileInputStream = new FileInputStream(file);
         BufferedInputStream bufferedInputStream = new BufferedInputStream(fileInputStream)) {
        
        // 读取文件所有字节到数组
        byte[] allBytes = bufferedInputStream.readAllBytes();
        // 过滤掉{和}对应的字节
        byte[] filteredBytes = new byte[allBytes.length];
        int index = 0;
        for (byte b : allBytes) {
            if (b != 0x7B && b != 0x7D) {
                filteredBytes[index++] = b;
            }
        }
        // 截取有效长度的数组
        filteredBytes = Arrays.copyOf(filteredBytes, index);
        
        CharsetDetector charsetDetector = new CharsetDetector();
        charsetDetector.setText(new ByteArrayInputStream(filteredBytes));
        charsetDetector.enableInputFilter(true);
        CharsetMatch cm = charsetDetector.detect();
        return cm != null ? cm.getName() : null;
    } catch (Exception e) {
        e.printStackTrace();
        return null;
    }
}

额外提醒

  • 上面的方案都是基于单字节编码的假设,如果你的文件是UTF-16这类双字节编码,{会对应两个字节(比如UTF-16LE是0x7B 0x00),这时候单字节过滤会出错。这种情况你可能需要先做一次粗略编码检测,再根据编码过滤对应的字节序列,但复杂度会高一些。不过一般来说,编码检测工具对少量干扰字符的容忍度还不错,大部分场景下方案1和2都够用。
  • 方案1适合大文件,不用加载全部字节到内存;方案2代码更简洁,适合小文件。

内容的提问来源于stack exchange,提问作者StackerSapper

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.07 18:02:33