JSON是否属于二进制编码格式?作为可读文本的它如何编码为字节序列?
Great question—let’s break this down clearly, since there’s a common confusion here between text-based formats and how they get converted to bytes for storage or transmission.
First, let’s clarify the core distinction:
- JSON is a text-based serialization format: Its entire structure is built from human-readable Unicode characters (like
{,",:, letters, numbers) that follow a strict syntax for representing data structures (objects, arrays, strings, numbers, booleans, null). - Binary encoding formats (like Protobuf, BSON, Thrift): These skip the human-readable text layer entirely—they encode data directly into raw bytes with no inherent character-based structure. You can’t open a Protobuf file in a text editor and make sense of it, but you can with JSON.
The confusion usually comes from two places:
- Every piece of data (text or binary) must be converted to a byte sequence to travel over a network or be written to a file. This step applies to JSON too, but it’s a separate process from the serialization format itself.
- In some workflows, JSON might be wrapped inside binary containers (like in certain message queues or database blobs), which can make it look like it’s a binary format at first glance. But the JSON itself is still text underneath.
The process has two distinct layers, which is easy to mix up:
Step 1: Serialize data to JSON text
First, your in-memory data (like a Python dictionary, Java object, etc.) gets converted into a valid JSON string—this is the human-readable text part. For example, a user object might become:
{"name": "Alice", "age": 30, "is_active": true}
Step 2: Encode the JSON text to bytes
Since computers only understand bytes, this Unicode text needs to be mapped to a byte sequence using a character encoding scheme. The JSON specification (RFC 8259) explicitly recommends UTF-8 as the preferred encoding (though UTF-16 and UTF-32 are allowed too).
Here’s how this works for a simple example:
- The character
{(Unicode code point U+007B) maps to the single byte0x7Bin UTF-8. - The string
"Alice"uses ASCII characters (all within U+0000 to U+007F), so each letter maps to one byte (e.g.,Ais0x41,lis0x6C). - If you have non-ASCII characters (like
{"name": "张三"}), the Chinese characters are encoded into multi-byte sequences: "张" (U+5F20) becomes0xE5 0xBC 0xA0, and "三" (U+4E09) becomes0xE4 0xB8 0x89in UTF-8.
When sending this over the network or writing to a file, you’re sending/writing that sequence of UTF-8 encoded bytes. On the receiving end, the bytes are first decoded back into Unicode text, then parsed into the original data structure.
内容的提问来源于stack exchange,提问作者user1008636

