如何将LZW压缩输出的整数数组转为字节并实现磁盘存读?
LZW整数编码与磁盘字节存储/读取实现
你的核心问题是将LZW压缩输出的整数数组转换为可存储的字节数据,再从磁盘读取还原为整数数组用于解压。下面直接给出基于Python struct 模块的实现方案:
核心思路
LZW生成的编码值会从256开始递增,单字节无法容纳,因此需要用多字节格式(如2字节或4字节无符号整数)存储这些整数。Python的struct模块可以方便地实现整数与字节串的双向转换。
完整实现代码
保留你原有的LZW压缩/解压函数,新增字节存储与读取的辅助函数:
import struct # 原有的LZW压缩函数 def compress(data): dictionary = {chr(i): i for i in range(256)} next_code = 256 result = [] sequence = "" for char in data: new_sequence = sequence + char if new_sequence in dictionary: sequence = new_sequence else: result.append(dictionary[sequence]) dictionary[new_sequence] = next_code next_code += 1 sequence = char if sequence: result.append(dictionary[sequence]) return result # 原有的LZW解压函数 def decompress(data): dictionary = {i: chr(i) for i in range(256)} next_code = 256 result = [] sequence = chr(data[0]) result.append(sequence) for code in data[1:]: if code in dictionary: entry = dictionary[code] elif code == next_code: entry = sequence + sequence[0] else: raise ValueError("Invalid compressed data") result.append(entry) dictionary[next_code] = sequence + entry[0] next_code += 1 sequence = entry return "".join(result) # 将LZW整数编码转为字节并保存到文件 def save_compressed(encoded_data, file_path, format=">H"): """ :param encoded_data: LZW压缩后的整数数组 :param file_path: 保存路径 :param format: struct打包格式,>H表示大端2字节无符号整数(最大支持65535),>I表示大端4字节无符号整数 """ # 批量打包整数为字节串 byte_data = struct.pack(f"{len(encoded_data)}{format}", *encoded_data) with open(file_path, "wb") as f: f.write(byte_data) # 从文件读取字节并还原为LZW整数编码数组 def load_compressed(file_path, format=">H"): """ :param file_path: 读取路径 :param format: 与保存时一致的struct格式 :return: LZW解压需要的整数数组 """ with open(file_path, "rb") as f: byte_data = f.read() # 计算整数个数,每个整数占用的字节数由format决定 code_size = struct.calcsize(format) if len(byte_data) % code_size != 0: raise ValueError("Corrupted compressed file: byte length mismatch") # 批量解包字节为整数数组 encoded_data = struct.unpack(f"{len(byte_data)//code_size}{format}", byte_data) return list(encoded_data)
使用示例
# 测试文本 test_data = "abracadabraabracadabra" # 压缩 encoded = compress(test_data) print("LZW编码数组:", encoded) # 保存到磁盘 save_compressed(encoded, "compressed.lzw") # 从磁盘读取 loaded_encoded = load_compressed("compressed.lzw") # 解压 decoded = decompress(loaded_encoded) print("解压结果:", decoded) print("是否与原数据一致:", decoded == test_data)
注意事项
- 格式选择:如果你的LZW字典可能超过65535(即编码值大于65535),请将
format参数改为>I(4字节无符号整数),避免溢出。 - 端序:
>表示大端字节序,跨平台存储时更推荐使用;如果不需要跨平台,也可以用<(小端)。 - 完整性校验:读取时检查字节长度是否为编码大小的整数倍,避免读取损坏的文件。
内容的提问来源于stack exchange,提问作者FoxMaccloud
相关产品推荐
相关产品推荐

