You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

JavaScript中String.fromCharCode转Buffer内容出现多余字符的原因

问题现象

在Node.js v16.14.2环境下,执行以下代码会得到带多余字符的结果:

> String.fromCharCode.apply(null, Buffer.from('°'))
'°'

但执行以下代码能得到正确结果:

> Buffer.from('°').toString('utf-8')
'°'

浏览器环境中使用ethereumjs.Buffer也出现相同问题:

console.log(String.fromCharCode.apply(null, ethereumjs.Buffer.Buffer.from('°'))) // '°'
console.log(ethereumjs.Buffer.Buffer.from('°').toString('utf-8')) // '°'
console.log(new Uint8Array(ethereumjs.Buffer.Buffer.from('°'))) // [ 194, 176 ]
console.log(new Uint8Array(ethereumjs.Buffer.Buffer.from(String.fromCharCode(176)))) // [ 194, 176 ]

同时发现特殊现象:

> new Uint8Array(Buffer.from('°'))
Uint8Array(2) [ 194, 176 ]
> new Uint8Array(Buffer.from(String.fromCharCode(176)))
Uint8Array(2) [ 194, 176 ]
原因解析

核心问题是**String.fromCharCode和Buffer.toString('utf-8')的处理逻辑完全不同**:

  • 字符°的UTF-8编码是两个字节:0xC2(十进制194)和0xB0(十进制176),这从Uint8Array的输出能直接看到。
  • String.fromCharCode是把每个传入的数字当作UCS-2/UTF-16代码单元转换,而非UTF-8字节。所以传入194和176时,会分别转成对应的Unicode字符:194对应Â(U+00C2),176对应°(U+00B0),拼起来就是°。
  • 而Buffer.toString('utf-8')是严格按照UTF-8编码规则,把字节序列解码成对应的Unicode字符,因此能得到正确的°。

至于那个特殊现象:Buffer.from(String.fromCharCode(176))得到[194,176],是因为String.fromCharCode(176)生成的是Unicode字符°(U+00B0),Buffer.from默认用UTF-8编码该字符,而U+00B0在UTF-8中恰好是两个字节0xC2 0xB0,所以结果和直接编码°一致。

内容的提问来源于stack exchange,提问作者Dipu

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.21 15:15:36