You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

UTF-8字节数组与字符串互转后数据不一致问题求助

问题描述

我遇到一个基础技术问题:我有一个存储double类型值22.22的8字节数组,其字节值为[184, 30, 133, 235, 81, 56, 54, 64]。我希望将它转换为字符串后,再还原为完全相同的字节数组,但转换后得到的是14字节的数组[239, 191, 189, 30, 239, 191, 189, 239, 191, 189, 81, 56, 54, 64],与原数组不一致。以下是我的代码:

List<byte[]> byteArrayList = new List<byte[]>();
...Here I populate the byteArrayList with 1 item - 8 bytes long
UTFEncoding encoding = new UTF8Encoding();
StringBuilder SB = new StringBuilder();
foreach(byte[] byteArrayItem in byteArrayList)//currently - for simplicity contains one item
{
    SB.Append(encoding.GetString(byteArrayItem));
}
string items = SB.ToString();
byte[] result = Encoding.UTF8.GetBytes(items); //I would expect the result to be identical to the byteArrayList  - it is not ! it contains 12 bytes !!!!
问题原因与解决方案

核心问题:误用了UTF-8编码的用途

UTF-8是文本编码,只能处理合法的UTF-8字符序列,不是用来承载任意二进制数据的。

  • 无效UTF-8字节会被替换:你原数组里的184、133、235这些字节,本身不是合法的UTF-8编码单元。当调用GetString()时,这些无效字节会被自动替换成Unicode的替换字符�(U+FFFD),这个字符用UTF-8编码就是3个字节[239,191,189]——这就是结果数组变长的原因。
  • 字符串不是二进制容器:字符串的设计目的是存储文本字符,不是直接保存二进制数据。强行把二进制转成字符串再转回,必然会因为编码规则篡改原始数据。

正确的解决方案:用Base64编码

Base64是专门为二进制→文本可逆转换设计的方案,能保证数据完全无损。修改后的代码如下:

List<byte[]> byteArrayList = new List<byte[]>();
// 填充8字节数组的代码
StringBuilder SB = new StringBuilder();
foreach(byte[] byteArrayItem in byteArrayList)
{
    // 把二进制数组转成Base64字符串
    SB.Append(Convert.ToBase64String(byteArrayItem));
}
string items = SB.ToString();

// 还原为原字节数组
byte[] result = Convert.FromBase64String(items);

这样处理后,还原得到的result会和原字节数组完全一致。

内容的提问来源于stack exchange,提问作者Guy E

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.04 03:10:12