如何从US-ASCII编码字节数组中正确显示西里尔字母?
解决西里尔字母转码后乱码的问题
你遇到的问题本质是US-ASCII编码不支持西里尔字符。US-ASCII仅包含0-127的字符,西里尔字母不在这个范围内,所以调用getBytes(StandardCharsets.US_ASCII)时,JVM会将无法编码的西里尔字符替换成?——这一步已经永久丢失了原始字符的信息,后续再转回字符串自然无法恢复。
要正确保留并还原西里尔字母,必须使用支持西里尔字符集的编码格式,推荐用UTF-8(通用跨平台)或Windows-1251(西里尔字符专用编码)。
方案1:使用UTF-8编码(推荐)
import java.nio.charset.StandardCharsets; import java.io.UnsupportedEncodingException; public class CyrillicEncodingTest { public static void main(String[] args) { // 使用UTF-8编码字符串为字节数组 byte[] bytes = "тестtest".getBytes(StandardCharsets.UTF_8); // 用UTF-8解码字节数组回字符串 System.out.println(getStringValue(bytes, StandardCharsets.UTF_8.name())); } private static String getStringValue(byte[] content, String encoding) { try { return new String(content, 0, content.length, encoding); } catch (UnsupportedEncodingException e) { e.printStackTrace(); return ""; } } }
方案2:使用Windows-1251编码(西里尔专用)
import java.nio.charset.Charset; import java.io.UnsupportedEncodingException; public class CyrillicEncodingTest { public static void main(String[] args) { // 使用Windows-1251编码字符串为字节数组 byte[] bytes = "тестtest".getBytes(Charset.forName("Windows-1251")); // 用Windows-1251解码字节数组回字符串 System.out.println(getStringValue(bytes, "Windows-1251")); } private static String getStringValue(byte[] content, String encoding) { try { return new String(content, 0, content.length, encoding); } catch (UnsupportedEncodingException e) { e.printStackTrace(); return ""; } } }
关键说明
- 两种方案都能正确输出
тестtest:UTF-8更适合跨平台场景,Windows-1251在仅处理西里尔字符时字节占用更小。 - 编码和解码必须使用同一套字符集,且该字符集要覆盖所有需要处理的字符,否则仍会出现字符丢失或乱码。
内容的提问来源于stack exchange,提问作者Костя Н
相关产品推荐
相关产品推荐

