为何Protobuf二进制文件体积比对应JSON文件更大?
我尝试用Protobuf加速前后端数据传输,做POC时加载JSON文件转成Protobuf缓冲区保存,结果新文件比原JSON还大。相关代码和结果如下:
定义的Protobuf文件(input.proto)
syntax = "proto3"; message MyData { repeated float a = 1; repeated float b = 2; repeated float c = 3; float d = 4; repeated float e = 5; }
序列化代码(index.mjs)
import protobuf from 'protobufjs'; import fs from 'fs'; protobuf.load('./input.proto', (err, root) => { const payload = JSON.parse(fs.readFileSync('./input.json', {encoding: 'utf8'})); var Message = root.lookupType("MyData"); var errMsg = Message.verify(payload); if (errMsg) throw Error(errMsg); var message = Message.create(payload); const buffer = Message.encode(message).finish(); fs.writeFileSync('./output.pb', buffer, 'binary'); }, () => { });
测试用JSON(input.json)
{ "a": [1, 2.4, 3, 4], "b": [1, 2, 3, 4], "c": [1, 2, 3.2, 4], "d": 10.321, "e": [1, 2, 3.7, 4] }
执行结果
$ du -h input.json output.pb 2,0M input.json 2,5M output.pb
(实际JSON文件比示例大得多,但格式一致)
问题原因分析
基础类型存储差异:
Protobuf的float是固定4字节二进制存储,而JSON里的数字是文本形式。比如JSON里的整数1只占1字节,转换成Protobuf的float要占4字节;即使是小数,比如10.321在JSON里是6字节文本,Protobuf的float是4字节,这种情况Protobuf更省,但如果数组里大量是短整数,JSON的文本形式总字节数会远低于Protobuf的二进制。重复字段的编码开销:
虽然Proto3中repeated基本类型默认是packed编码(整个数组作为连续块,仅带一次字段标签),但如果你的protobufjs版本或配置有问题,可能会退化成非packed模式——每个数组元素都带字段号和类型标识,这会额外增加大量字节开销。
解决办法
确保启用Packed编码
显式在proto字段中声明[packed=true],避免潜在的编码模式问题:syntax = "proto3"; message MyData { repeated float a = 1 [packed=true]; repeated float b = 2 [packed=true]; repeated float c = 3 [packed=true]; float d = 4; repeated float e = 5 [packed=true]; }优化字段类型
如果数组中存在大量整数,可以拆分字段类型,用int32存储整数(Protobuf的int32用varint编码,整数越小占用字节越少),单独用float存储小数,减少不必要的字节开销。对Protobuf输出进行压缩
Protobuf本身是二进制格式,压缩率极高。可以在序列化后用gzip等算法压缩,对比压缩后的体积:// 修改序列化代码,添加gzip压缩 import zlib from 'zlib'; // ... const buffer = Message.encode(message).finish(); const compressedBuffer = zlib.gzipSync(buffer); fs.writeFileSync('./output.pb.gz', compressedBuffer, 'binary');压缩后的Protobuf体积通常会远小于压缩后的JSON。
内容的提问来源于stack exchange,提问作者Zorzi

