使用Apache Commons IO的IOUtils.copy转换MultipartFile为String时Unicode字符丢失怎么办?
Let’s break down why your Unicode characters are getting mangled (like Ble痊脕no turning into Ble?脕no) and fix this step by step.
First, the most common culprit is the uploaded file isn’t actually saved in UTF-8, even if you expect it to be. If users save the file in another encoding (like GBK, GB2312, or Windows-1252), forcing UTF-8 parsing in your code will replace unrecognizable bytes with ? characters.
Start with a quick check: Open the original file in a tool like Notepad++ or VS Code, and verify its actual encoding (look for the "Encoding" menu in Notepad++). If the file isn’t UTF-8, either ask users to save it as UTF-8, or add logic to detect and handle the file’s encoding (though auto-detection is tricky—standardizing on UTF-8 is better).
If the file is definitely UTF-8, let’s refine your code to ensure no encoding slips happen during reading:
1. Explicitly Specify Encoding with InputStreamReader
Your existing code uses IOUtils.copy(stream, writer, Charsets.UTF_8), but wrapping the input stream in an InputStreamReader with a clear UTF-8 specification eliminates any implicit encoding ambiguity:
import java.io.InputStream; import java.io.InputStreamReader; import java.io.StringWriter; import java.nio.charset.StandardCharsets; import org.apache.commons.io.IOUtils; public String getUploadFileAsString() { // Use try-with-resources to auto-close streams and avoid leaks try (final InputStream stream = file.getInputStream(); final InputStreamReader reader = new InputStreamReader(stream, StandardCharsets.UTF_8); final StringWriter writer = new StringWriter()) { IOUtils.copy(reader, writer); return writer.toString(); } catch (final IOException e) { throw new RuntimeException("Failed to read uploaded file content", e); } }
2. Handle UTF-8 BOM (Byte Order Mark)
Some UTF-8 files include a BOM (a hidden EF BB BF byte sequence at the start). This can confuse parsers and lead to unexpected character issues. Use Apache Commons IO’s BOMInputStream to skip the BOM safely:
import org.apache.commons.io.input.BOMInputStream; import java.nio.charset.StandardCharsets; import static org.apache.commons.io.ByteOrderMark.UTF_8; public String getUploadFileAsString() { try (final InputStream rawStream = file.getInputStream(); // Skip UTF-8 BOM while preserving the rest of the content final InputStream stream = new BOMInputStream(rawStream, false, UTF_8); final InputStreamReader reader = new InputStreamReader(stream, StandardCharsets.UTF_8); final StringWriter writer = new StringWriter()) { IOUtils.copy(reader, writer); return writer.toString(); } catch (final IOException e) { throw new RuntimeException("Failed to read uploaded file content", e); } }
3. Check Post-Reading Processing
If reading works but characters still get lost later, verify these downstream steps:
- Ensure your database tables use
utf8mb4(not justutf8) to support full Unicode, including emojis and rare characters. - Set your API response’s
Content-Typeto includecharset=UTF-8(e.g.,application/json; charset=UTF-8). - Avoid using
String.getBytes()without specifying an encoding—always usegetBytes(StandardCharsets.UTF_8)instead of relying on the system default.
By following these steps, you’ll ensure Unicode characters are preserved from upload through to final processing.
内容的提问来源于stack exchange,提问作者Naveen Kumar Gautam

