You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Apache Commons IO的IOUtils.copy转换MultipartFile为String时Unicode字符丢失怎么办?

Fixing Unicode Character Loss When Uploading UTF-8 Text Files

Let’s break down why your Unicode characters are getting mangled (like Ble痊脕no turning into Ble?脕no) and fix this step by step.

First, the most common culprit is the uploaded file isn’t actually saved in UTF-8, even if you expect it to be. If users save the file in another encoding (like GBK, GB2312, or Windows-1252), forcing UTF-8 parsing in your code will replace unrecognizable bytes with ? characters.

Start with a quick check: Open the original file in a tool like Notepad++ or VS Code, and verify its actual encoding (look for the "Encoding" menu in Notepad++). If the file isn’t UTF-8, either ask users to save it as UTF-8, or add logic to detect and handle the file’s encoding (though auto-detection is tricky—standardizing on UTF-8 is better).

If the file is definitely UTF-8, let’s refine your code to ensure no encoding slips happen during reading:

1. Explicitly Specify Encoding with InputStreamReader

Your existing code uses IOUtils.copy(stream, writer, Charsets.UTF_8), but wrapping the input stream in an InputStreamReader with a clear UTF-8 specification eliminates any implicit encoding ambiguity:

import java.io.InputStream;
import java.io.InputStreamReader;
import java.io.StringWriter;
import java.nio.charset.StandardCharsets;
import org.apache.commons.io.IOUtils;

public String getUploadFileAsString() {
    // Use try-with-resources to auto-close streams and avoid leaks
    try (final InputStream stream = file.getInputStream();
         final InputStreamReader reader = new InputStreamReader(stream, StandardCharsets.UTF_8);
         final StringWriter writer = new StringWriter()) {
        IOUtils.copy(reader, writer);
        return writer.toString();
    } catch (final IOException e) {
        throw new RuntimeException("Failed to read uploaded file content", e);
    }
}

2. Handle UTF-8 BOM (Byte Order Mark)

Some UTF-8 files include a BOM (a hidden EF BB BF byte sequence at the start). This can confuse parsers and lead to unexpected character issues. Use Apache Commons IO’s BOMInputStream to skip the BOM safely:

import org.apache.commons.io.input.BOMInputStream;
import java.nio.charset.StandardCharsets;
import static org.apache.commons.io.ByteOrderMark.UTF_8;

public String getUploadFileAsString() {
    try (final InputStream rawStream = file.getInputStream();
         // Skip UTF-8 BOM while preserving the rest of the content
         final InputStream stream = new BOMInputStream(rawStream, false, UTF_8);
         final InputStreamReader reader = new InputStreamReader(stream, StandardCharsets.UTF_8);
         final StringWriter writer = new StringWriter()) {
        IOUtils.copy(reader, writer);
        return writer.toString();
    } catch (final IOException e) {
        throw new RuntimeException("Failed to read uploaded file content", e);
    }
}

3. Check Post-Reading Processing

If reading works but characters still get lost later, verify these downstream steps:

  • Ensure your database tables use utf8mb4 (not just utf8) to support full Unicode, including emojis and rare characters.
  • Set your API response’s Content-Type to include charset=UTF-8 (e.g., application/json; charset=UTF-8).
  • Avoid using String.getBytes() without specifying an encoding—always use getBytes(StandardCharsets.UTF_8) instead of relying on the system default.

By following these steps, you’ll ensure Unicode characters are preserved from upload through to final processing.

内容的提问来源于stack exchange,提问作者Naveen Kumar Gautam

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 06:36:29