Java如何替换InputStream中的字符串并写入文件
问题原因
你的代码不生效主要是两个核心问题:
- 搜索关键词定义错误:你写的
"\"https\"".getBytes("UTF-8")实际匹配的是带前后双引号的"https"字符串,如果你要替换的是不带引号的https,搜索词从根上就写错了,自然不会触发替换。 - 自定义
ReplacingInputStream存在逻辑缺陷:- 把代表流结束的
-1标记也存入了inQueue队列,会打乱匹配判断逻辑 - 没有处理流读取到末尾后,
inQueue里剩余的、长度不足搜索词长度的残留内容,容易出现内容截断 - 部分JDK的
Files.copy实现会直接调用批量字节读取方法,如果你没有重写对应方法,自定义的单字节读取替换逻辑不会被触发。
- 把代表流结束的
推荐实现方案
你不强制要求流式替换的话,没必要自己实现复杂的流匹配逻辑,直接按字符集读取流内容、完成替换后写入文件即可,代码简单不容易出错:
// 注意指定和源内容一致的字符集,不要用默认字符集避免乱码,这里以UTF-8为例 String content = new String(inputStream.readAllBytes(), StandardCharsets.UTF_8); // 执行替换,比如要移除所有https content = content.replace("https", ""); // 写入目标文件 Files.writeString(file.toPath(), content, StandardCharsets.UTF_8, StandardCopyOption.REPLACE_EXISTING);
Java 8版本没有
readAllBytes()/writeString()方法,可以用ByteArrayOutputStream中转读取:ByteArrayOutputStream result = new ByteArrayOutputStream(); byte[] buffer = new byte[1024]; int length; while ((length = inputStream.read(buffer)) != -1) { result.write(buffer, 0, length); } String content = result.toString(StandardCharsets.UTF_8.name()); // 替换逻辑和上面一致,写完内容后用FileOutputStream写入文件即可
流式替换修正方案
如果你确实需要InputStream到InputStream的流式替换(比如处理GB级大文件不想全量加载到内存),可以用修正后的ReplacingInputStream实现,重写批量读取方法,同时修复匹配逻辑的bug:
class ReplacingInputStream extends FilterInputStream { private final LinkedList<Integer> inQueue = new LinkedList<>(); private final LinkedList<Integer> outQueue = new LinkedList<>(); private final byte[] search; private final byte[] replacement; private boolean eofReached = false; protected ReplacingInputStream(InputStream in, byte[] search, byte[] replacement) { super(in); this.search = search; this.replacement = replacement; } private boolean isMatchFound() { if (inQueue.size() < search.length) return false; Iterator<Integer> iter = inQueue.iterator(); for (byte b : search) { if (b != iter.next()) return false; } return true; } private void readAhead() throws IOException { while (!eofReached && inQueue.size() < search.length) { int next = super.read(); if (next == -1) { eofReached = true; break; } inQueue.offer(next); } } @Override public int read() throws IOException { if (!outQueue.isEmpty()) { return outQueue.poll(); } readAhead(); if (isMatchFound()) { // 匹配到搜索词,移除匹配的字节,写入替换内容 for (int i = 0; i < search.length; i++) { inQueue.poll(); } for (byte b : replacement) { outQueue.offer((int) b); } return read(); } // 没匹配到,输出队列第一个字节;已经到流末尾且队列为空时返回-1 if (!inQueue.isEmpty()) { return inQueue.poll(); } return eofReached ? -1 : read(); } @Override public int read(byte[] b, int off, int len) throws IOException { // 重写批量读取方法,避免默认实现的性能问题/逻辑不生效问题 int readCount = 0; for (int i = 0; i < len; i++) { int next = read(); if (next == -1) { return readCount == 0 ? -1 : readCount; } b[off + i] = (byte) next; readCount++; } return readCount; } }
调用时注意搜索词不要写错,如果要移除不带引号的https,搜索词定义改为:
byte[] search = "https".getBytes(StandardCharsets.UTF_8); byte[] replacement = "".getBytes(StandardCharsets.UTF_8);
内容的提问来源于stack exchange,提问作者kiki keke
相关产品推荐
相关产品推荐

