Java中如何检查字符所属指定字符集及识别转码失败字符
Let's walk through fixing your method to properly detect and collect characters that can't be transcoded, especially for the UTF-8 to Latin1 case.
First, let's address a small bug in your current code: the line String originalLine = new String(outputData); is incorrect because it uses the JVM's default charset to decode the output bytes (which are in toCharset). This line doesn't serve any purpose here, so we can remove it entirely.
Next, the core issue: to capture the original characters that can't be converted, we need to check each character from the decoded input (not the converted output, which only has replacement characters like ?). Here's how to implement the correct check:
- Extract the original characters from the decoded
CharBuffer(this is the actual text we started with after decoding the input string). - Create a
CharsetEncoderfor the target charset once (more efficient than creating it per character). - For each original character, check if the encoder can convert it without replacement. If not, add it to your
notEncodedbuilder.
Here's the corrected code with these changes:
private String transcodeLineFromTo(String string, Charset fromCharset, Charset toCharset) { try { ByteBuffer inputBuffer = ByteBuffer.wrap(string.getBytes(fromCharset)); CharBuffer data = fromCharset.decode(inputBuffer); ByteBuffer outputBuffer = toCharset.encode(data); byte[] outputData = outputBuffer.array(); String convertedLine = new String(outputData, toCharset); StringBuilder notEncoded = new StringBuilder(); CharsetEncoder encoder = toCharset.newEncoder(); char[] originalCharacters = data.toString().toCharArray(); for (char ch : originalCharacters) { // Check if the character can't be encoded by the target charset if (!encoder.canEncode(ch)) { notEncoded.append(ch); } } // Follow your requirement: return unconvertible chars if any (for UTF-8 to Latin1), else converted line return notEncoded.length() > 0 ? notEncoded.toString() : convertedLine; } catch (Exception e) { throw new IllegalStateException(e); } }
Key Notes:
- UTF-8 to Latin1: Since Latin1 only supports characters in the 0-255 range, any UTF-8 character outside this set will be flagged and collected in
notEncoded. The method will return these unconvertible characters if they exist, which matches your requirement. - Latin1 to UTF-8: All Latin1 characters are fully representable in UTF-8, so
notEncodedwill always be empty, and the method returns the converted line as expected. - Efficiency: Creating the
CharsetEncoderonce outside the loop avoids redundant object creation, which is better for performance, especially if this method is called frequently. - Replacement Characters: The default encoder behavior replaces unconvertible characters with
?, but our check catches the original characters before this replacement happens, which is exactly what you need.
内容的提问来源于stack exchange,提问作者Alexandra Dediu

