Huffman编码图像压缩异常咨询:压缩后文件体积增大
图像Huffman压缩异常问题解答
一、关于原图像大小的疑问
1. 原图像大小是否正确?
你看到的318KB是JPG格式图像的磁盘存储大小,但JPG本身是经过有损压缩的格式,而你的代码处理的是图像解码后的原始像素数据(RGB图像每个像素占24位,即3字节),两者不是同一维度的大小。比如假设你的图像是1024×1024分辨率的RGB图,原始像素数据大小是1024×1024×3=3145728字节≈3MB,远大于JPG的318KB,这是正常现象。
2. PC中如何获取图像的原始像素数据大小?
- 手动计算:用图像工具(比如Photoshop、系统自带的画图3D)查看图像的宽、高和通道数,通过公式
宽×高×通道数得到字节数,再除以1024转换为KB。 - 代码获取:用你现有代码里的
image = np.asarray(Image.open(path)),然后执行print(image.nbytes)就能直接得到原始像素数据的字节大小,除以1024就是对应的KB数。
二、压缩后文件暴增的核心问题
你的代码存在两个关键错误,直接导致压缩后体积反超原文件:
1. 将二进制编码以文本字符存储
Huffman编码的结果是一串二进制比特流(比如010110),但你直接把这串字符串写入了文本文件compressed.txt。文本文件中每个'0'或'1'字符会占用1个字节(ASCII编码),相当于把原本1位的信息膨胀成了8位,体积直接放大8倍。
比如原本压缩后只需要100万位,写成文本后就变成100万字节(≈976KB),如果用二进制存储的话仅需125KB(100万÷8)。
2. 未保存Huffman解码树
你的代码只保存了编码后的字符串,但没有保存Huffman树的结构。没有解码树,根本无法将编码内容还原为原始图像,这个压缩文件本质是无效的。
三、修复建议
1. 改用二进制方式存储编码数据
不要直接写入字符串,把编码的比特流转换成字节后再写入二进制文件:
# 将编码字符串转为字节 encoding_str = encoding # 补全到8的倍数,方便后续解码 padding = 8 - len(encoding_str) % 8 encoding_str += '0' * padding # 按每8位一组转换为字节数组 byte_array = bytearray([int(encoding_str[i:i+8], 2) for i in range(0, len(encoding_str), 8)]) # 写入二进制文件,同时记录补零的数量 with open('compressed.bin', 'wb') as f: f.write(padding.to_bytes(1, byteorder='big')) f.write(byte_array)
2. 保存Huffman树结构
需要把Huffman树序列化后单独保存,比如用pickle模块:
import pickle with open('huffman_tree.pkl', 'wb') as f: pickle.dump(tree, f)
后续解码时,再反序列化读取树结构即可。
3. 正确对比压缩效果
对比的基准应该是原始像素数据的大小(而非JPG的压缩后大小)和压缩后的二进制文件+树文件的总大小。Huffman编码对已经经过JPG压缩的像素数据(冗余已被大量去除)压缩效果有限,但不会出现体积膨胀的情况。
附:修正缩进后的原代码
path = input('Enter image path:') image = np.asarray(Image.open(path)) pixels = [] for row in image: for ch in row: for pix in ch: pixels.append(pix) class Node: def __init__(self, prob, symbol, left=None, right=None): # probability of symbol self.prob = prob # symbol self.symbol = symbol # left node self.left = left # right node self.right = right # tree direction (0/1) self.code = '' """ A function to print the codes of symbols by traveling Huffman Tree""" codes = dict() def Calculate_Codes(node, val=''): # huffman code for current node newVal = val + str(node.code) if(node.left): Calculate_Codes(node.left, newVal) if(node.right): Calculate_Codes(node.right, newVal) if(not node.left and not node.right): codes[node.symbol] = newVal return codes """ A function to get the probabilities of symbols in given data""" def Calculate_Probability(data): symbols = dict() for element in data: if symbols.get(element) == None: symbols[element] = 1 else: symbols[element] += 1 return symbols """ A function to obtain the encoded output""" def Output_Encoded(data, coding): encoding_output = [] for c in data: # print(coding[c], end = '') encoding_output.append(coding[c]) string = ''.join([str(item) for item in encoding_output]) return string """ A function to calculate the space difference between compressed and non compressed data""" def Total_Gain(data, coding): before_compression = len(data) * 8 # total bit space to store the data before compression after_compression = 0 symbols = coding.keys() for symbol in symbols: count = data.count(symbol) after_compression += count * len(coding[symbol]) #calculate how many bit is required for that symbol in total print("Space usage before compression (in bits):", before_compression) print("Space usage after compression (in bits):", after_compression) def Huffman_Encoding(data): symbol_with_probs = Calculate_Probability(data) symbols = symbol_with_probs.keys() probabilities = symbol_with_probs.values() print("symbols: ", symbols) print("probabilities: ", probabilities) nodes = [] # converting symbols and probabilities into huffman tree nodes for symbol in symbols: nodes.append(Node(symbol_with_probs.get(symbol), symbol)) while len(nodes) > 1: # sort all the nodes in ascending order based on their probability nodes = sorted(nodes, key=lambda x: x.prob) # for node in nodes: # print(node.symbol, node.prob) # pick 2 smallest nodes right = nodes[0] left = nodes[1] left.code = 0 right.code = 1 # combine the 2 smallest nodes to create new node newNode = Node(left.prob+right.prob, left.symbol+right.symbol, left, right) nodes.remove(left) nodes.remove(right) nodes.append(newNode) huffman_encoding = Calculate_Codes(nodes[0]) print("symbols with codes", huffman_encoding) Total_Gain(data, huffman_encoding) encoded_output = Output_Encoded(data,huffman_encoding) return encoded_output, nodes[0] encoding, tree = Huffman_Encoding(pixels) file = open('compressed.txt','w') file.write(encoding) file.close()
内容的提问来源于stack exchange,提问作者hamid
相关产品推荐
相关产品推荐

