You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

SpringFramework HTTP请求Unicode字符乱码问题求助

解决Spring HTTP请求非UTF-8编码页面乱码问题

问题背景

通过Spring Framework实现了HTTP请求工具方法,核心代码如下:

public <T>HttpReply<T> httpRequest(final String uri, final HttpMethod method,
            final Class<T> expectedReturnType, final List<HttpMessageConverter<?>> messageConverters,
            final HashMap<String, Object> formValues, final HashMap<String, Object> headers)
                    throws HttpNullUriOrMethodException, HttpInvocationException {
        try {

            redirectInfo.set(new AbstractMap.SimpleEntry<String, String>(uri, ""));

            if (method==null) {
                throw new HttpNullUriOrMethodException("HttpMethod cannot be null.");
            }

            if (!StringUtils.hasText(uri)) {
                throw new HttpNullUriOrMethodException("URI cannot be null or empty.");
            }

            HttpRequestExecutingMessageHandler handler =
                    buildMessageHandler(uri, method, expectedReturnType, messageConverters);

            // Default queue for reply
            QueueChannel replyChannel = new QueueChannel();
            handler.setOutputChannel(replyChannel);

            // Exec Http Request
            Message<?> message = buildMessage(formValues, headers);
            try {
                handler.handleMessage(message);
            }
            catch (Exception e) {
                throw new HttpInvocationException("Error Handling HTTP Message.");
            }

            // Get Response
            Message<?> response = replyChannel.receive();
            if (response == null) {
                throw new HttpInvocationException("Error: communication is interrupted.");
            }

            // Read response Headers
            String[] usefulHeaders = readUsefulHeaders(response.getHeaders());

            // Return payload
            Object respObj = response.getPayload();             

            if (expectedReturnType != null && !expectedReturnType.isInstance(respObj)) {
                throw new HttpInvocationException("Error: response payload is instance of "
                         + respObj.getClass().getName() + ". Expected: " + expectedReturnType.getClass().getName());
            }

            HttpReply<T> retVal = new HttpReply<>();
            retVal.setPayload((T)respObj);

            String valRedirect = uri;
            if (redirectInfo.get().getKey().equals(uri)) {
                if (StringUtils.hasText(redirectInfo.get().getValue())) {
                    valRedirect = redirectInfo.get().getValue();
                }
            }
            else {
                throw new HttpInvocationException("ERROR READING REDIRECT INFORMATION!!! Original URI: "
                        + uri + " - FOUND URI: " + redirectInfo.get().getKey());
            }
            retVal.setActualLocation(valRedirect);
            return retVal;
        }
        finally {
            redirectInfo.remove();
        }
    }

调用方式:

HttpReply<byte[]> feedContent = httpUtil.httpRequest(rssFeed.getUrl(), HttpMethod.GET, byte[].class, null,
                null, null);

rawXml = new String(feedContent.getPayload());

运行过程中,读取非UTF-8编码页面时,rawXml频繁出现�乱码。尝试过硬编码设置handler.setCharset(StandardCharsets.ISO_8859_1)、修改请求头contentType=application/xml; charset=ISO-8859-1,或获取字节后手动转码,但遇到GBK等其他编码时,无法解决乱码问题。

解决方案:自动识别响应编码

乱码的核心原因是硬编码指定字符集,或未从响应头/内容中自动识别正确编码。以下是具体修复步骤:

1. 从响应头Content-Type提取编码

在获取响应后,先解析Content-Type头中的charset参数:

// 在获取response后添加编码识别逻辑
HttpHeaders responseHeaders = (HttpHeaders) response.getHeaders();
String contentType = responseHeaders.getFirst(HttpHeaders.CONTENT_TYPE);
String charset = null;
if (contentType != null) {
    MediaType mediaType = MediaType.parseMediaType(contentType);
    if (mediaType.getCharset() != null) {
        charset = mediaType.getCharset().name();
    }
}

2. 从XML声明中提取编码(针对RSS/XML场景)

如果响应头未指定编码,XML文档通常会在开头声明中指定编码(如<?xml version="1.0" encoding="GBK"?>),可通过字节流解析该声明:

// 若响应头无编码,尝试从XML声明中提取
if (charset == null && respObj instanceof byte[]) {
    byte[] payload = (byte[]) respObj;
    // 取前100字节解析XML声明,用ISO-8859-1保证不破坏原始字节
    String xmlDeclPrefix = new String(payload, 0, Math.min(100, payload.length), StandardCharsets.ISO_8859_1);
    Pattern encodingPattern = Pattern.compile("encoding=\"([^\"]+)\"");
    Matcher matcher = encodingPattern.matcher(xmlDeclPrefix);
    if (matcher.find()) {
        charset = matcher.group(1);
    }
}

3. 设置兜底编码

若以上步骤都未找到编码,设置兜底编码(如UTF-8)避免完全乱码:

if (charset == null) {
    charset = StandardCharsets.UTF_8.name();
}

4. 完善HttpReply并正确转换字符串

修改HttpReply类,新增charset字段存储识别到的编码:

// 示例:HttpReply类新增字段和getter/setter
public class HttpReply<T> {
    private T payload;
    private String actualLocation;
    private String charset; // 新增字段

    // getter和setter方法
    public String getCharset() { return charset; }
    public void setCharset(String charset) { this.charset = charset; }
}

在工具方法中设置该字段:

retVal.setCharset(charset);

调用时使用正确编码转换字节数组:

rawXml = new String(feedContent.getPayload(), feedContent.getCharset());

5. 移除硬编码的handler charset设置

删除handler.setCharset(...)这类硬编码配置,让Spring保持原始字节流,我们在最后一步根据识别到的编码转换即可。

内容的提问来源于stack exchange,提问作者Malignus

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.06 18:30:49