如何移除BufferedInputStream中的{}字符以进行文件编码检测?
如何移除BufferedInputStream中的指定字符后用于编码检测?
我懂你的痛点——字节流可不像字符串那样有现成的replace方法,而且咱们要先处理流再检测编码,还不能提前把字节转成字符(不然就预设了编码,那检测还有啥意义对吧?)。这里有两个实用方案,你可以根据文件大小来选:
方案1:自定义过滤InputStream(适合大文件)
我们可以继承FilterInputStream,重写读取方法,直接跳过{和}对应的字节。这里要说明下:在UTF-8、GBK、ISO-8859-1这类常见编码里,{对应的字节是0x7B,}是0x7D,都是单字节,所以过滤这两个值就能覆盖绝大多数场景。
先写一个自定义的过滤流类:
import java.io.FilterInputStream; import java.io.IOException; import java.io.InputStream; public class CharFilterInputStream extends FilterInputStream { // 定义{和}对应的字节值 private static final byte LEFT_BRACE = 0x7B; private static final byte RIGHT_BRACE = 0x7D; protected CharFilterInputStream(InputStream in) { super(in); } // 重写单个字节读取方法,跳过目标字符 @Override public int read() throws IOException { int b; do { b = super.read(); } while (b == LEFT_BRACE || b == RIGHT_BRACE); return b; } // 重写批量读取方法,过滤掉目标字节后返回 @Override public int read(byte[] b, int off, int len) throws IOException { int bytesRead = super.read(b, off, len); if (bytesRead == -1) { return -1; } int newIndex = off; for (int i = off; i < off + bytesRead; i++) { if (b[i] != LEFT_BRACE && b[i] != RIGHT_BRACE) { b[newIndex++] = b[i]; } } return newIndex - off; } }
然后修改你的编码检测方法,把原有的BufferedInputStream包装到这个自定义流里就行:
static String detectCharset(File file) { try (FileInputStream fileInputStream = new FileInputStream(file); BufferedInputStream bufferedInputStream = new BufferedInputStream(fileInputStream); CharFilterInputStream filteredInputStream = new CharFilterInputStream(bufferedInputStream)) { CharsetDetector charsetDetector = new CharsetDetector(); charsetDetector.setText(filteredInputStream); charsetDetector.enableInputFilter(true); CharsetMatch cm = charsetDetector.detect(); return cm != null ? cm.getName() : null; } catch (Exception e) { e.printStackTrace(); return null; } }
方案2:读取到字节数组后过滤(适合小文件)
如果你的文件不大,直接把所有字节读到内存里过滤会更简单,代码量也少:
import java.util.Arrays; static String detectCharset(File file) { try (FileInputStream fileInputStream = new FileInputStream(file); BufferedInputStream bufferedInputStream = new BufferedInputStream(fileInputStream)) { // 读取文件所有字节到数组 byte[] allBytes = bufferedInputStream.readAllBytes(); // 过滤掉{和}对应的字节 byte[] filteredBytes = new byte[allBytes.length]; int index = 0; for (byte b : allBytes) { if (b != 0x7B && b != 0x7D) { filteredBytes[index++] = b; } } // 截取有效长度的数组 filteredBytes = Arrays.copyOf(filteredBytes, index); CharsetDetector charsetDetector = new CharsetDetector(); charsetDetector.setText(new ByteArrayInputStream(filteredBytes)); charsetDetector.enableInputFilter(true); CharsetMatch cm = charsetDetector.detect(); return cm != null ? cm.getName() : null; } catch (Exception e) { e.printStackTrace(); return null; } }
额外提醒
- 上面的方案都是基于单字节编码的假设,如果你的文件是UTF-16这类双字节编码,
{会对应两个字节(比如UTF-16LE是0x7B 0x00),这时候单字节过滤会出错。这种情况你可能需要先做一次粗略编码检测,再根据编码过滤对应的字节序列,但复杂度会高一些。不过一般来说,编码检测工具对少量干扰字符的容忍度还不错,大部分场景下方案1和2都够用。 - 方案1适合大文件,不用加载全部字节到内存;方案2代码更简洁,适合小文件。
内容的提问来源于stack exchange,提问作者StackerSapper
相关产品推荐
相关产品推荐

