Java HTTP代理响应体解析的编码与读取超时问题求解
解决HTTP代理的编码与响应体读取超时问题
问题背景
开发一款仅支持HTTP的代理,核心逻辑是:接收客户端GET请求→转发至目标主机→修改目标返回的HTML响应体→调整响应头(如Content-Length)后返回给客户端。
当前遇到两个核心问题:
- 编码乱码:返回给浏览器的HTML出现大量“带问号的菱形”字符;
- 读取超时:读取响应体时无法读到Content-Length指定的字节数,导致循环超时(例如179643字节的响应体仅读到3000字节就停滞,触发5秒延迟)。
问题代码片段
问题出在响应体读取的循环逻辑中:
private Response getResponse(final Socket socket) { try { HashMap<String, String> headers; StringBuilder builder = new StringBuilder(); BufferedReader stream = new BufferedReader(new InputStreamReader(socket.getInputStream())); //-- READ FIRST LINE --// // We assume that it is a valid response! (TODO) String[] firstLine = stream.readLine().split(" ",3); headers = getHeaders(stream); //--- GET BODY --- String contentLength = headers.get("Content-Length"); //Check if body exists if(contentLength != null) { int bodyLength = Integer.parseInt(contentLength); String s; //The issue occurs in this while loop! while(bodyLength > 0 && (s = stream.readLine()) != null) { bodyLength -= (s+"\n").getBytes(StandardCharsets.UTF_8).length; builder.append(s).append("\n"); } } //-- Return Request --- int code = Integer.parseInt(firstLine[1]); return new Response(headers,builder, firstLine[0],code, firstLine[2]); } catch (IOException e) { e.printStackTrace(); return null; } }
核心疑问
Java的String内部采用UTF-16编码,使用InputStreamReader(socket.getInputStream(), "UTF-8")读取时,原响应体的编码会丢失吗?读取的字符串是UTF-8编码的,还是直接转成了UTF-16的String对象?
已尝试方案
- 给
InputStreamReader指定UTF-8编码,输出流同步设置,解决了乱码但超时问题依旧; - 尝试按字节解析响应体,失败且不利于后续内容修改。
问题根源与解决方案
1. 超时问题根源与修复
根源:
readLine()依赖换行符分割内容,但HTTP响应体不一定是按行组织的(比如长文本、二进制内容),没有换行符时readLine()会一直阻塞等待,直到超时;- 字节数计算逻辑错误:用
(s+"\n").getBytes(UTF_8).length计算的是字符串转UTF-8后的字节数,和原响应体的原始字节数不匹配(比如原响应体是GBK编码,转UTF-8后字节数会增加),导致bodyLength永远无法减到0,循环持续阻塞。
修复方案:
放弃按行读取,直接读取原始字节流,严格按Content-Length指定的字节数读取:
private Response getResponse(final Socket socket) { try { HashMap<String, String> headers; StringBuilder bodyBuilder = new StringBuilder(); // 仅用BufferedReader读取响应头(文本格式) BufferedReader headerReader = new BufferedReader(new InputStreamReader(socket.getInputStream())); String[] firstLine = headerReader.readLine().split(" ", 3); headers = getHeaders(headerReader); // 获取原始字节流读取响应体 InputStream inputStream = socket.getInputStream(); String contentLength = headers.get("Content-Length"); if (contentLength != null) { int totalBytes = Integer.parseInt(contentLength); byte[] bodyBytes = new byte[totalBytes]; int bytesRead = 0; // 循环读取直到读满指定字节数 while (bytesRead < totalBytes) { int readCount = inputStream.read(bodyBytes, bytesRead, totalBytes - bytesRead); if (readCount == -1) { throw new IOException("Stream closed before reading full body"); } bytesRead += readCount; } // 从Content-Type头部提取编码,无指定则用默认UTF-8 String charset = "UTF-8"; String contentType = headers.get("Content-Type"); if (contentType != null && contentType.contains("charset=")) { charset = contentType.split("charset=")[1].split(";")[0].trim(); } // 转成字符串用于修改 String body = new String(bodyBytes, charset); // 此处添加内容修改逻辑:body = modifyHtmlContent(body); bodyBuilder.append(body); // 后续返回响应时,需将修改后的body转成对应编码的字节数组,更新Content-Length头部 } int code = Integer.parseInt(firstLine[1]); return new Response(headers, bodyBuilder, firstLine[0], code, firstLine[2]); } catch (IOException e) { e.printStackTrace(); return null; } }
2. 编码问题根源与修复
根源:
Java的String内部是UTF-16编码,InputStreamReader的作用是将原始字节流按指定编码解码为Unicode字符(存储为UTF-16的String),原响应体的编码信息不会被String保存。如果硬编码解码方式(比如固定UTF-8),会导致非UTF-8编码的页面解码错误,出现乱码。
修复方案:
- 从响应头
Content-Type中提取charset参数(例如Content-Type: text/html; charset=GBK),用该编码解码原始字节; - 修改内容后,用相同编码将字符串转回字节数组,同时更新响应头的
Content-Length为新字节数组的长度。
注意事项
- 不要混用
BufferedReader和直接读取InputStream:BufferedReader会缓冲部分字节,导致直接读取InputStream时丢失数据; - 必须严格按
Content-Length的字节数读取响应体,避免依赖换行符等文本格式特征; - 编码处理需动态适配响应头指定的字符集,不能硬编码。
内容的提问来源于stack exchange,提问作者Dubstepzedd
相关产品推荐
相关产品推荐

