You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将LZW压缩输出的整数数组转为字节并实现磁盘存读?

LZW整数编码与磁盘字节存储/读取实现

你的核心问题是将LZW压缩输出的整数数组转换为可存储的字节数据,再从磁盘读取还原为整数数组用于解压。下面直接给出基于Python struct 模块的实现方案:

核心思路

LZW生成的编码值会从256开始递增,单字节无法容纳,因此需要用多字节格式(如2字节或4字节无符号整数)存储这些整数。Python的struct模块可以方便地实现整数与字节串的双向转换。

完整实现代码

保留你原有的LZW压缩/解压函数,新增字节存储与读取的辅助函数:

import struct

# 原有的LZW压缩函数
def compress(data):
    dictionary = {chr(i): i for i in range(256)}
    next_code = 256
    result = []
    sequence = ""

    for char in data:
        new_sequence = sequence + char

        if new_sequence in dictionary:
            sequence = new_sequence
        else:
            result.append(dictionary[sequence])
            dictionary[new_sequence] = next_code
            next_code += 1
            sequence = char

    if sequence:
        result.append(dictionary[sequence])

    return result

# 原有的LZW解压函数
def decompress(data):
    dictionary = {i: chr(i) for i in range(256)}
    next_code = 256
    result = []

    sequence = chr(data[0])
    result.append(sequence)

    for code in data[1:]:
        if code in dictionary:
            entry = dictionary[code]
        elif code == next_code:
            entry = sequence + sequence[0]
        else:
            raise ValueError("Invalid compressed data")

        result.append(entry)
        dictionary[next_code] = sequence + entry[0]
        next_code += 1
        sequence = entry

    return "".join(result)

# 将LZW整数编码转为字节并保存到文件
def save_compressed(encoded_data, file_path, format=">H"):
    """
    :param encoded_data: LZW压缩后的整数数组
    :param file_path: 保存路径
    :param format: struct打包格式,>H表示大端2字节无符号整数(最大支持65535),>I表示大端4字节无符号整数
    """
    # 批量打包整数为字节串
    byte_data = struct.pack(f"{len(encoded_data)}{format}", *encoded_data)
    with open(file_path, "wb") as f:
        f.write(byte_data)

# 从文件读取字节并还原为LZW整数编码数组
def load_compressed(file_path, format=">H"):
    """
    :param file_path: 读取路径
    :param format: 与保存时一致的struct格式
    :return: LZW解压需要的整数数组
    """
    with open(file_path, "rb") as f:
        byte_data = f.read()
    
    # 计算整数个数,每个整数占用的字节数由format决定
    code_size = struct.calcsize(format)
    if len(byte_data) % code_size != 0:
        raise ValueError("Corrupted compressed file: byte length mismatch")
    
    # 批量解包字节为整数数组
    encoded_data = struct.unpack(f"{len(byte_data)//code_size}{format}", byte_data)
    return list(encoded_data)

使用示例

# 测试文本
test_data = "abracadabraabracadabra"

# 压缩
encoded = compress(test_data)
print("LZW编码数组:", encoded)

# 保存到磁盘
save_compressed(encoded, "compressed.lzw")

# 从磁盘读取
loaded_encoded = load_compressed("compressed.lzw")

# 解压
decoded = decompress(loaded_encoded)
print("解压结果:", decoded)
print("是否与原数据一致:", decoded == test_data)

注意事项

  • 格式选择:如果你的LZW字典可能超过65535(即编码值大于65535),请将format参数改为>I(4字节无符号整数),避免溢出。
  • 端序:>表示大端字节序,跨平台存储时更推荐使用;如果不需要跨平台,也可以用<(小端)。
  • 完整性校验:读取时检查字节长度是否为编码大小的整数倍,避免读取损坏的文件。

内容的提问来源于stack exchange,提问作者FoxMaccloud

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.09 22:51:00