You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Java HTTP代理响应体解析的编码与读取超时问题求解

解决HTTP代理的编码与响应体读取超时问题

问题背景

开发一款仅支持HTTP的代理,核心逻辑是:接收客户端GET请求→转发至目标主机→修改目标返回的HTML响应体→调整响应头(如Content-Length)后返回给客户端。

当前遇到两个核心问题:

  1. 编码乱码:返回给浏览器的HTML出现大量“带问号的菱形”字符;
  2. 读取超时:读取响应体时无法读到Content-Length指定的字节数,导致循环超时(例如179643字节的响应体仅读到3000字节就停滞,触发5秒延迟)。

问题代码片段

问题出在响应体读取的循环逻辑中:

private Response getResponse(final Socket socket) {
    try {

        HashMap<String, String> headers;
        StringBuilder builder = new StringBuilder();
        BufferedReader stream = new BufferedReader(new InputStreamReader(socket.getInputStream()));

        //-- READ FIRST LINE --//
        // We assume that it is a valid response! (TODO)
        String[] firstLine = stream.readLine().split(" ",3);
        headers = getHeaders(stream);

        //--- GET BODY ---
        String contentLength = headers.get("Content-Length");
        //Check if body exists
        if(contentLength != null) {
            int bodyLength = Integer.parseInt(contentLength);

            String s;
            //The issue occurs in this while loop!
            while(bodyLength > 0 && (s = stream.readLine()) != null) {
                bodyLength -= (s+"\n").getBytes(StandardCharsets.UTF_8).length;
                builder.append(s).append("\n");
            }

        }

        //-- Return Request ---
        int code = Integer.parseInt(firstLine[1]);
        return new Response(headers,builder, firstLine[0],code, firstLine[2]);
    }
    catch (IOException e) {
        e.printStackTrace();
        return null;
    }
}

核心疑问

Java的String内部采用UTF-16编码,使用InputStreamReader(socket.getInputStream(), "UTF-8")读取时,原响应体的编码会丢失吗?读取的字符串是UTF-8编码的,还是直接转成了UTF-16的String对象?

已尝试方案

  • 给InputStreamReader指定UTF-8编码,输出流同步设置,解决了乱码但超时问题依旧;
  • 尝试按字节解析响应体,失败且不利于后续内容修改。

问题根源与解决方案

1. 超时问题根源与修复

根源:

  • readLine()依赖换行符分割内容,但HTTP响应体不一定是按行组织的(比如长文本、二进制内容),没有换行符时readLine()会一直阻塞等待,直到超时;
  • 字节数计算逻辑错误:用(s+"\n").getBytes(UTF_8).length计算的是字符串转UTF-8后的字节数,和原响应体的原始字节数不匹配(比如原响应体是GBK编码,转UTF-8后字节数会增加),导致bodyLength永远无法减到0,循环持续阻塞。

修复方案:
放弃按行读取,直接读取原始字节流,严格按Content-Length指定的字节数读取:

private Response getResponse(final Socket socket) {
    try {
        HashMap<String, String> headers;
        StringBuilder bodyBuilder = new StringBuilder();
        
        // 仅用BufferedReader读取响应头(文本格式)
        BufferedReader headerReader = new BufferedReader(new InputStreamReader(socket.getInputStream()));
        String[] firstLine = headerReader.readLine().split(" ", 3);
        headers = getHeaders(headerReader);

        // 获取原始字节流读取响应体
        InputStream inputStream = socket.getInputStream();
        String contentLength = headers.get("Content-Length");
        
        if (contentLength != null) {
            int totalBytes = Integer.parseInt(contentLength);
            byte[] bodyBytes = new byte[totalBytes];
            int bytesRead = 0;
            
            // 循环读取直到读满指定字节数
            while (bytesRead < totalBytes) {
                int readCount = inputStream.read(bodyBytes, bytesRead, totalBytes - bytesRead);
                if (readCount == -1) {
                    throw new IOException("Stream closed before reading full body");
                }
                bytesRead += readCount;
            }

            // 从Content-Type头部提取编码,无指定则用默认UTF-8
            String charset = "UTF-8";
            String contentType = headers.get("Content-Type");
            if (contentType != null && contentType.contains("charset=")) {
                charset = contentType.split("charset=")[1].split(";")[0].trim();
            }

            // 转成字符串用于修改
            String body = new String(bodyBytes, charset);
            // 此处添加内容修改逻辑:body = modifyHtmlContent(body);
            
            bodyBuilder.append(body);
            // 后续返回响应时,需将修改后的body转成对应编码的字节数组,更新Content-Length头部
        }

        int code = Integer.parseInt(firstLine[1]);
        return new Response(headers, bodyBuilder, firstLine[0], code, firstLine[2]);
    } catch (IOException e) {
        e.printStackTrace();
        return null;
    }
}

2. 编码问题根源与修复

根源:
Java的String内部是UTF-16编码,InputStreamReader的作用是将原始字节流按指定编码解码为Unicode字符(存储为UTF-16的String),原响应体的编码信息不会被String保存。如果硬编码解码方式(比如固定UTF-8),会导致非UTF-8编码的页面解码错误,出现乱码。

修复方案:

  • 从响应头Content-Type中提取charset参数(例如Content-Type: text/html; charset=GBK),用该编码解码原始字节;
  • 修改内容后,用相同编码将字符串转回字节数组,同时更新响应头的Content-Length为新字节数组的长度。

注意事项

  • 不要混用BufferedReader和直接读取InputStream:BufferedReader会缓冲部分字节,导致直接读取InputStream时丢失数据;
  • 必须严格按Content-Length的字节数读取响应体,避免依赖换行符等文本格式特征;
  • 编码处理需动态适配响应头指定的字符集,不能硬编码。

内容的提问来源于stack exchange,提问作者Dubstepzedd

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.26 02:47:04