You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

学习Python OpenCV时对bytearray函数UTF-16编码结果的困惑求解

Understanding bytearray() output with UTF-16 encoding in Python

Let's break down what's happening here step by step, since your confusion comes from how UTF-16 encoding works (including byte order and the BOM):

1. First, clarify the structure of your output

Your output is:

bytearray(b'\xff\xfeG\x00e\x00e\x00k\x00s\x00f\x00o\x00r\x00g\x00e\x00e\x00k\x00s\x00')

You mentioned seeing \xfeG, but that's a visual mix-up—\xff\xfe is a separate prefix (the Byte Order Mark, BOM), and the first character's bytes start right after that: G\x00.

2. What is \xff\xfe?

UTF-16 encodes every Unicode character into 2 or 4 bytes, but byte order (which byte comes first) can vary between systems:

  • Big Endian (BE): Higher byte first
  • Little Endian (LE): Lower byte first

To tell decoders which order to use, UTF-16 uses a BOM at the start of the byte stream:

  • \xff\xfe = Little Endian (LE) BOM
  • \xfe\xff = Big Endian (BE) BOM

When you use 'utf-16' as the encoding in Python, it defaults to adding the LE BOM (which is why you see \xff\xfe at the start).

3. Why do you see G\x00 for the character 'G'?

The uppercase 'G' has a Unicode code point of U+0047. For ASCII-range characters (U+0000 to U+00FF), UTF-16 uses two bytes: the higher byte is 0x00, and the lower byte is the ASCII value of the character.

Since we're using Little Endian (LE) order, the lower byte comes first:

  • Lower byte: 0x47 (which is the ASCII byte for 'G')
  • Higher byte: 0x00 (null byte, represented as \x00)

So the two bytes for 'G' are 0x47 0x00, which shows up as G\x00 in the output (because 0x47 is printable as 'G').

4. Verify with other characters

Let's check 'e' (Unicode U+0065):

  • Lower byte: 0x65 (ASCII 'e')
  • Higher byte: 0x00
  • LE UTF-16 bytes: 0x65 0x00 → shows up as e\x00 in your output, which matches perfectly.

5. What if you use Big Endian?

If you specify 'utf-16-be' (no BOM by default), the output would look like this:

bytearray(b'\x00G\x00e\x00e\x00k\x00s\x00f\x00o\x00r\x00g\x00e\x00e\x00k\x00s')

Here, the higher byte (0x00) comes first for each character, which is the reverse of your original output.

Hopefully this clears up why the bytearray looks the way it does—you're just seeing UTF-16's Little Endian encoding with its required BOM!

内容的提问来源于stack exchange,提问作者Alex

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.29 18:12:40