You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

整数列表二进制表示的文件压缩解压缩:解包字节长度标识与高效压缩方案咨询

哈哈,这个问题我之前做二进制序列化的时候踩过坑!核心矛盾就是「不定长数据怎么让解码端知道边界」,既要尽量少占空间,又要解码高效。下面给你几个从优到次的方案,你可以根据自己的场景挑:

最优方案:可变长度整数编码(Varint)

这是兼顾文件体积和解析效率的首选,也是Protobuf、Flatbuffers等序列化框架的核心方案之一。

原理

每个字节的最高位作为「续位标记」:如果该位是1,说明后面还有字节属于当前整数;如果是0,说明这是最后一个字节。剩下的7位用来存储整数的二进制数据(小端序逻辑)。这样小整数(比如0-127)只占1字节,大整数会多占几个字节,但平均下来体积比固定长度小很多,解码时也能边读边解析,效率很高。

代码实现

编码函数

def varint_encode(n):
    bytes_list = []
    while True:
        # 取低7位数据
        byte = n & 0x7F
        # 右移7位,处理剩下的部分
        n = n >> 7
        if n > 0:
            # 如果还有后续字节,把最高位设为1
            bytes_list.append(byte | 0x80)
        else:
            bytes_list.append(byte)
            break
    return bytes(bytes_list)

解码函数

def varint_decode(stream):
    result = 0
    shift = 0
    while True:
        byte_data = stream.read(1)
        if not byte_data:
            raise EOFError("Unexpected end of stream")
        byte = ord(byte_data)
        # 把当前字节的7位数据拼到结果里
        result |= (byte & 0x7F) << shift
        # 如果最高位是0,说明结束了
        if not (byte & 0x80):
            break
        shift += 7
    return result

完整读写示例

def write_varints(outFileName, int_list):
    with open(outFileName, "wb") as f:
        for num in int_list:
            f.write(varint_encode(num))

def read_varints(inFileName):
    int_list = []
    with open(inFileName, "rb") as f:
        while True:
            try:
                num = varint_decode(f)
                int_list.append(num)
            except EOFError:
                break
    return int_list

优缺点

  • ✅ 体积最优:小整数仅占1字节,远优于固定长度方案
  • ✅ 解码高效:无需额外存储长度信息,边读边解析
  • ❌ 超大整数(比如超过64位)需要额外处理,但绝大多数场景下够用

次优方案:带Varint长度前缀的不定长字节

如果你的整数长度差异极大,或者不想修改整数本身的编码逻辑,可以在每个整数的字节数据前,用Varint编码存储该整数的字节长度。这样解码端先读长度,再读对应字节数的数据即可。

代码实现

编码

import math

def encode_with_varint_len(outFileName, int_list):
    with open(outFileName, "wb") as f:
        for num in int_list:
            num_bytes = math.ceil(num.bit_length() / 8)
            # 先写长度的Varint编码
            f.write(varint_encode(num_bytes))
            # 再写整数本身的字节数据
            f.write(num.to_bytes(num_bytes, 'little'))

解码

def decode_with_varint_len(inFileName):
    int_list = []
    with open(inFileName, "rb") as f:
        while True:
            try:
                # 先读字节长度
                num_bytes = varint_decode(f)
                # 再读对应长度的字节数据
                byte_data = f.read(num_bytes)
                if not byte_data:
                    break
                num = int.from_bytes(byte_data, 'little')
                int_list.append(num)
            except EOFError:
                break
    return int_list

优缺点

  • ✅ 逻辑简单,兼容你原来的to_bytes编码逻辑
  • ✅ 长度前缀本身占用空间小(最多2字节就能表示0-16383的长度,完全够用)
  • ❌ 比纯Varint编码多了长度前缀的开销,体积略大

备选方案:固定长度前缀

如果你的场景更看重解码速度,对体积要求不高,可以用固定字节数存储每个整数的长度(比如2字节,能覆盖0-65535的字节长度,足够容纳64位整数)。

代码实现

编码

def encode_with_fixed_len(outFileName, int_list):
    with open(outFileName, "wb") as f:
        for num in int_list:
            num_bytes = math.ceil(num.bit_length() / 8)
            # 用2字节固定长度存储字节数(小端序)
            f.write(num_bytes.to_bytes(2, 'little'))
            f.write(num.to_bytes(num_bytes, 'little'))

解码

def decode_with_fixed_len(inFileName):
    int_list = []
    with open(inFileName, "rb") as f:
        while True:
            # 直接读2字节的长度
            len_data = f.read(2)
            if not len_data:
                break
            num_bytes = int.from_bytes(len_data, 'little')
            byte_data = f.read(num_bytes)
            num = int.from_bytes(byte_data, 'little')
            int_list.append(num)
    return int_list

优缺点

  • ✅ 解码速度最快:不需要循环解析长度,直接读固定字节即可
  • ❌ 体积最大:即使小整数,也要多写2字节的长度前缀

总结

如果想兼顾最小文件体积和高效解析,Varint编码绝对是最优选择,这也是工业界的通用方案。如果你的整数大多是小数值,这个方案能把文件体积压到极致,同时解码速度也不会慢。

内容的提问来源于stack exchange,提问作者b20

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.29 03:54:06