You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Node.js与.NET Framework 4.6.1的UTF-8编码差异咨询

UTF-8编码转换:Node.js与.NET Framework 4.6.1的差异解决

问题场景

有一个字节数组 [65, 119, 212, 250, 152, 244, 166],使用UTF-8编码转换为字符串时,Node.js与.NET Framework 4.6.1得到了不同的结果:

Node.js 代码示例

// Node.js
const b7 = Buffer.from([65, 119, 212, 250, 152, 244, 166 ]).toString();
console.log(b7.length === 7); // true
console.log(b7 === 'Aw�����'); // true

执行后得到长度为7的字符串。

.NET Framework 4.6.1 代码示例

// C# .NET
byte[] b6 = new byte[7] { 65, 119, 212, 250, 152, 244, 166 };
UTF8Encoding encoding = new UTF8Encoding();
string s6 = encoding.GetString(b6);
Console.WriteLine(s6.Length == 6); // true
Console.WriteLine(s6 == "Aw����"); // true

执行后得到长度为6的字符串。

问题解答

1. 为什么.NET版本的字符串长度更短?

核心原因是两者对无效UTF-8序列的错误处理逻辑不同:

  • 给定的字节数组包含多处无效UTF-8序列:
    • 212+250:212是双字节UTF-8的起始字节,但后续的250不属于合法的后续字节范围(0x80-0xBF),构成无效序列。
    • 152:单独的后续字节,无对应起始字节,属于无效序列。
    • 244+166:244是四字节UTF-8的起始字节,但缺少必要的两个后续字节,构成无效序列。
  • .NET的UTF8Encoding.GetString默认会将连续的无效字节序列替换为单个U+FFFD(替换字符):
    • 212+250 → 1个U+FFFD
    • 152 → 1个U+FFFD
    • 244+166 → 1个U+FFFD
      加上前两个合法字符A和w,总长度为2+3=6。
  • Node.js默认的Buffer.toString('utf8')会将每个无效字节单独替换为U+FFFD:
    • 212、250、152、244、166每个字节都对应一个U+FFFD
      加上前两个合法字符,总长度为2+5=7。

2. 如何使Node.js的转换结果与.NET一致?

要让Node.js的处理逻辑和.NET对齐,需要使用TextDecoderAPI,它对无效UTF-8序列的处理规则和.NET一致——将连续的无效序列合并为单个替换字符。

示例代码:

const decoder = new TextDecoder('utf-8');
const bytes = new Uint8Array([65, 119, 212, 250, 152, 244, 166]);
const s = decoder.decode(bytes);
console.log(s.length === 6); // true
console.log(s === 'Aw����'); // true

如果需要基于Buffer处理,可先将Buffer转换为Uint8Array再传入TextDecoder,效果完全一致。


内容的提问来源于stack exchange,提问作者zolipapa

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.14 21:55:31