You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Java中如何检查字符所属指定字符集及识别转码失败字符

Fixing Character Transcoding Status Check for UTF-8 ↔ Latin1

Let's walk through fixing your method to properly detect and collect characters that can't be transcoded, especially for the UTF-8 to Latin1 case.

First, let's address a small bug in your current code: the line String originalLine = new String(outputData); is incorrect because it uses the JVM's default charset to decode the output bytes (which are in toCharset). This line doesn't serve any purpose here, so we can remove it entirely.

Next, the core issue: to capture the original characters that can't be converted, we need to check each character from the decoded input (not the converted output, which only has replacement characters like ?). Here's how to implement the correct check:

  1. Extract the original characters from the decoded CharBuffer (this is the actual text we started with after decoding the input string).
  2. Create a CharsetEncoder for the target charset once (more efficient than creating it per character).
  3. For each original character, check if the encoder can convert it without replacement. If not, add it to your notEncoded builder.

Here's the corrected code with these changes:

private String transcodeLineFromTo(String string, Charset fromCharset, Charset toCharset) {
    try {
        ByteBuffer inputBuffer = ByteBuffer.wrap(string.getBytes(fromCharset));
        CharBuffer data = fromCharset.decode(inputBuffer);
        ByteBuffer outputBuffer = toCharset.encode(data);
        byte[] outputData = outputBuffer.array();
        String convertedLine = new String(outputData, toCharset);
        
        StringBuilder notEncoded = new StringBuilder();
        CharsetEncoder encoder = toCharset.newEncoder();
        char[] originalCharacters = data.toString().toCharArray();
        
        for (char ch : originalCharacters) {
            // Check if the character can't be encoded by the target charset
            if (!encoder.canEncode(ch)) {
                notEncoded.append(ch);
            }
        }
        
        // Follow your requirement: return unconvertible chars if any (for UTF-8 to Latin1), else converted line
        return notEncoded.length() > 0 ? notEncoded.toString() : convertedLine;
    } catch (Exception e) {
        throw new IllegalStateException(e);
    }
}

Key Notes:

  • UTF-8 to Latin1: Since Latin1 only supports characters in the 0-255 range, any UTF-8 character outside this set will be flagged and collected in notEncoded. The method will return these unconvertible characters if they exist, which matches your requirement.
  • Latin1 to UTF-8: All Latin1 characters are fully representable in UTF-8, so notEncoded will always be empty, and the method returns the converted line as expected.
  • Efficiency: Creating the CharsetEncoder once outside the loop avoids redundant object creation, which is better for performance, especially if this method is called frequently.
  • Replacement Characters: The default encoder behavior replaces unconvertible characters with ?, but our check catches the original characters before this replacement happens, which is exactly what you need.

内容的提问来源于stack exchange,提问作者Alexandra Dediu

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.07 14:48:12