UTF-8字节数组与字符串互转后数据不一致问题求助
问题描述
我遇到一个基础技术问题:我有一个存储double类型值22.22的8字节数组,其字节值为[184, 30, 133, 235, 81, 56, 54, 64]。我希望将它转换为字符串后,再还原为完全相同的字节数组,但转换后得到的是14字节的数组[239, 191, 189, 30, 239, 191, 189, 239, 191, 189, 81, 56, 54, 64],与原数组不一致。以下是我的代码:
List<byte[]> byteArrayList = new List<byte[]>(); ...Here I populate the byteArrayList with 1 item - 8 bytes long UTFEncoding encoding = new UTF8Encoding(); StringBuilder SB = new StringBuilder(); foreach(byte[] byteArrayItem in byteArrayList)//currently - for simplicity contains one item { SB.Append(encoding.GetString(byteArrayItem)); } string items = SB.ToString(); byte[] result = Encoding.UTF8.GetBytes(items); //I would expect the result to be identical to the byteArrayList - it is not ! it contains 12 bytes !!!!
问题原因与解决方案
核心问题:误用了UTF-8编码的用途
UTF-8是文本编码,只能处理合法的UTF-8字符序列,不是用来承载任意二进制数据的。
- 无效UTF-8字节会被替换:你原数组里的
184、133、235这些字节,本身不是合法的UTF-8编码单元。当调用GetString()时,这些无效字节会被自动替换成Unicode的替换字符�(U+FFFD),这个字符用UTF-8编码就是3个字节[239,191,189]——这就是结果数组变长的原因。 - 字符串不是二进制容器:字符串的设计目的是存储文本字符,不是直接保存二进制数据。强行把二进制转成字符串再转回,必然会因为编码规则篡改原始数据。
正确的解决方案:用Base64编码
Base64是专门为二进制→文本可逆转换设计的方案,能保证数据完全无损。修改后的代码如下:
List<byte[]> byteArrayList = new List<byte[]>(); // 填充8字节数组的代码 StringBuilder SB = new StringBuilder(); foreach(byte[] byteArrayItem in byteArrayList) { // 把二进制数组转成Base64字符串 SB.Append(Convert.ToBase64String(byteArrayItem)); } string items = SB.ToString(); // 还原为原字节数组 byte[] result = Convert.FromBase64String(items);
这样处理后,还原得到的result会和原字节数组完全一致。
内容的提问来源于stack exchange,提问作者Guy E
相关产品推荐
相关产品推荐

