You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

C#中从流中检测\r\n\r\n字节序列的最快方法

处理流读取直到\r\n\r\n的可靠方案

嘿,这个场景我熟!要从流里读字节直到碰到\r\n\r\n(也就是字节序列[13,10,13,10]),确实不能傻呵呵按4字节块去查——毕竟这序列很可能跨在两次读取的块中间,比如前一次读了3个字节\r\n\r,下一次读的第一个字节刚好是\n,凑成完整序列。下面给你两种实用的实现思路,附代码示例,都是经过实际验证的:

思路1:逐个读取+滑动窗口检查

这种方法的核心是维护一个长度为4的“滑动窗口”,每次读一个字节就更新窗口,然后检查是否匹配目标序列。不管序列出现在哪个位置,都能精准捕捉到。

代码示例(Java版本)

import java.io.InputStream;
import java.io.IOException;
import java.util.ArrayList;
import java.util.List;

public class StreamReader {
    // 定义终止序列:\r\n\r\n
    private static final byte[] TERMINATOR = {13, 10, 13, 10};

    public static byte[] readUntilTerminator(InputStream inputStream) throws IOException {
        List<Byte> contentBuffer = new ArrayList<>();
        byte[] slidingWindow = new byte[4];
        int currentByte;

        while ((currentByte = inputStream.read()) != -1) {
            byte byteVal = (byte) currentByte;
            contentBuffer.add(byteVal);

            // 更新滑动窗口:把前三个字节后移,新字节补到最后
            slidingWindow[0] = slidingWindow[1];
            slidingWindow[1] = slidingWindow[2];
            slidingWindow[2] = slidingWindow[3];
            slidingWindow[3] = byteVal;

            // 检查窗口是否匹配终止序列
            if (isTerminatorMatch(slidingWindow)) {
                // 移除终止序列本身,返回前面的内容
                contentBuffer.remove(contentBuffer.size() - 4);
                break;
            }
        }

        // 把List转成byte数组返回
        byte[] result = new byte[contentBuffer.size()];
        for (int i = 0; i < contentBuffer.size(); i++) {
            result[i] = contentBuffer.get(i);
        }
        return result;
    }

    private static boolean isTerminatorMatch(byte[] window) {
        for (int i = 0; i < TERMINATOR.length; i++) {
            if (window[i] != TERMINATOR[i]) {
                return false;
            }
        }
        return true;
    }
}

思路解析

  • 用ArrayList动态存已读字节,不用提前预估缓冲区大小,灵活得很
  • 滑动窗口始终盯着最后4个字节,哪怕终止序列是跨了多次读取凑成的,只要最后4个字节匹配,立刻就能检测到
  • 找到终止序列后,直接把这4个字节从缓冲区里删掉,返回前面的有效内容,完全符合需求

思路2:批量读取+滑动窗口优化

如果流的数据量很大,逐个读效率太低,就可以改成批量读取,同时处理跨块的终止序列。核心思路是每次读一块字节,然后在块里找终止序列;如果没找到,就把块的最后3个字节保留下来(因为终止序列是4字节,最多只需要前一个块的最后3个,和新块拼接就能覆盖跨块情况),下次读取时拼接再查。

代码示例(Java版本)

import java.io.InputStream;
import java.io.IOException;
import java.util.Arrays;

public class BatchStreamReader {
    private static final byte[] TERMINATOR = {13, 10, 13, 10};
    private static final int BATCH_SIZE = 1024; // 每次批量读1024字节,可按需调整

    public static byte[] readUntilTerminator(InputStream inputStream) throws IOException {
        byte[] batchBuffer = new byte[BATCH_SIZE];
        byte[] leftoverBytes = new byte[0]; // 存储上一次读取的末尾最多3字节
        int totalRead = 0;

        while (true) {
            int bytesRead = inputStream.read(batchBuffer);
            if (bytesRead == -1) {
                // 流提前结束,返回已读的所有内容
                byte[] result = new byte[leftoverBytes.length + totalRead];
                System.arraycopy(leftoverBytes, 0, result, 0, leftoverBytes.length);
                System.arraycopy(batchBuffer, 0, result, leftoverBytes.length, totalRead);
                return result;
            }

            // 拼接上一次的剩余字节和当前读取的块
            byte[] combinedData = new byte[leftoverBytes.length + bytesRead];
            System.arraycopy(leftoverBytes, 0, combinedData, 0, leftoverBytes.length);
            System.arraycopy(batchBuffer, 0, combinedData, leftoverBytes.length, bytesRead);

            // 在拼接后的数组里找终止序列的位置
            int terminatorPos = findTerminatorPosition(combinedData);
            if (terminatorPos != -1) {
                // 找到终止序列,返回序列之前的内容
                return Arrays.copyOf(combinedData, terminatorPos);
            } else {
                // 没找到,更新剩余字节和总读取量
                totalRead += bytesRead;
                // 最多保留最后3字节,避免冗余
                leftoverBytes = combinedData.length >= 3 
                    ? Arrays.copyOfRange(combinedData, combinedData.length - 3, combinedData.length) 
                    : combinedData;
            }
        }
    }

    private static int findTerminatorPosition(byte[] data) {
        for (int i = 0; i <= data.length - TERMINATOR.length; i++) {
            boolean match = true;
            for (int j = 0; j < TERMINATOR.length; j++) {
                if (data[i + j] != TERMINATOR[j]) {
                    match = false;
                    break;
                }
            }
            if (match) {
                return i;
            }
        }
        return -1;
    }
}

思路解析

  • 批量读取大幅提升效率,适合大流量场景,批量大小可以根据实际情况调整
  • leftoverBytes完美解决了跨块检测的问题,确保不会漏掉任何可能的终止序列
  • 提前处理了流提前结束的情况,就算没读到终止序列,也能返回已经读取的所有内容

关键注意事项

  • 绝对不能直接按4字节块读取并检查!比如前一个块的最后2字节是\r\n,下一个块的前2字节是\r\n,按4字节块读就会完全错过这个终止序列
  • 滑动窗口是核心逻辑,不管是逐个读还是批量读,都需要维护可能和终止序列相关的末尾字节
  • 一定要考虑流提前结束的情况,避免出现数据丢失或者异常

内容的提问来源于stack exchange,提问作者kmeshavkin

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 06:47:13