You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Huffman编码图像压缩异常咨询:压缩后文件体积增大

图像Huffman压缩异常问题解答

一、关于原图像大小的疑问

1. 原图像大小是否正确?

你看到的318KB是JPG格式图像的磁盘存储大小,但JPG本身是经过有损压缩的格式,而你的代码处理的是图像解码后的原始像素数据(RGB图像每个像素占24位,即3字节),两者不是同一维度的大小。比如假设你的图像是1024×1024分辨率的RGB图,原始像素数据大小是1024×1024×3=3145728字节≈3MB,远大于JPG的318KB,这是正常现象。

2. PC中如何获取图像的原始像素数据大小?

  • 手动计算:用图像工具(比如Photoshop、系统自带的画图3D)查看图像的宽、高和通道数,通过公式 宽×高×通道数 得到字节数,再除以1024转换为KB。
  • 代码获取:用你现有代码里的image = np.asarray(Image.open(path)),然后执行print(image.nbytes)就能直接得到原始像素数据的字节大小,除以1024就是对应的KB数。

二、压缩后文件暴增的核心问题

你的代码存在两个关键错误,直接导致压缩后体积反超原文件:

1. 将二进制编码以文本字符存储

Huffman编码的结果是一串二进制比特流(比如010110),但你直接把这串字符串写入了文本文件compressed.txt。文本文件中每个'0'或'1'字符会占用1个字节(ASCII编码),相当于把原本1位的信息膨胀成了8位,体积直接放大8倍。

比如原本压缩后只需要100万位,写成文本后就变成100万字节(≈976KB),如果用二进制存储的话仅需125KB(100万÷8)。

2. 未保存Huffman解码树

你的代码只保存了编码后的字符串,但没有保存Huffman树的结构。没有解码树,根本无法将编码内容还原为原始图像,这个压缩文件本质是无效的。

三、修复建议

1. 改用二进制方式存储编码数据

不要直接写入字符串,把编码的比特流转换成字节后再写入二进制文件:

# 将编码字符串转为字节
encoding_str = encoding
# 补全到8的倍数,方便后续解码
padding = 8 - len(encoding_str) % 8
encoding_str += '0' * padding
# 按每8位一组转换为字节数组
byte_array = bytearray([int(encoding_str[i:i+8], 2) for i in range(0, len(encoding_str), 8)])
# 写入二进制文件,同时记录补零的数量
with open('compressed.bin', 'wb') as f:
    f.write(padding.to_bytes(1, byteorder='big'))
    f.write(byte_array)

2. 保存Huffman树结构

需要把Huffman树序列化后单独保存,比如用pickle模块:

import pickle
with open('huffman_tree.pkl', 'wb') as f:
    pickle.dump(tree, f)

后续解码时,再反序列化读取树结构即可。

3. 正确对比压缩效果

对比的基准应该是原始像素数据的大小(而非JPG的压缩后大小)和压缩后的二进制文件+树文件的总大小。Huffman编码对已经经过JPG压缩的像素数据(冗余已被大量去除)压缩效果有限,但不会出现体积膨胀的情况。

附:修正缩进后的原代码

path = input('Enter image path:')

image = np.asarray(Image.open(path))

pixels = []

for row in image:
    for ch in row:
        for pix in ch:
            pixels.append(pix)

class Node:
    def __init__(self, prob, symbol, left=None, right=None):
        # probability of symbol
        self.prob = prob
        # symbol 
        self.symbol = symbol
        # left node
        self.left = left
        # right node
        self.right = right
        # tree direction (0/1)
        self.code = ''

""" A function to print the codes of symbols by traveling Huffman Tree"""
codes = dict()

def Calculate_Codes(node, val=''):
    # huffman code for current node
    newVal = val + str(node.code)
    if(node.left):
       Calculate_Codes(node.left, newVal)
    if(node.right):
       Calculate_Codes(node.right, newVal)
    if(not node.left and not node.right):
       codes[node.symbol] = newVal
    return codes        

""" A  function to get the probabilities of symbols in given data"""
def Calculate_Probability(data):
    symbols = dict()
    for element in data:
        if symbols.get(element) == None:
           symbols[element] = 1
        else: 
           symbols[element] += 1     
    return symbols

""" A function to obtain the encoded output"""
def Output_Encoded(data, coding):
    encoding_output = []
    for c in data:
        #  print(coding[c], end = '')
        encoding_output.append(coding[c])
    string = ''.join([str(item) for item in encoding_output])    
    return string
        
""" A function to calculate the space difference between compressed and non compressed data"""    
def Total_Gain(data, coding):
     before_compression = len(data) * 8 # total bit space to store the data before compression
     after_compression = 0
     symbols = coding.keys()
     for symbol in symbols:
         count = data.count(symbol)
         after_compression += count * len(coding[symbol]) #calculate how many bit is required for that symbol in total
     print("Space usage before compression (in bits):", before_compression)    
     print("Space usage after compression (in bits):", after_compression)           

def Huffman_Encoding(data):
    symbol_with_probs = Calculate_Probability(data)
    symbols = symbol_with_probs.keys()
    probabilities = symbol_with_probs.values()
    print("symbols: ", symbols)
    print("probabilities: ", probabilities)
    
    nodes = []
    
    # converting symbols and probabilities into huffman tree nodes
    for symbol in symbols:
        nodes.append(Node(symbol_with_probs.get(symbol), symbol))
    
    while len(nodes) > 1:
          # sort all the nodes in ascending order based on their probability
          nodes = sorted(nodes, key=lambda x: x.prob)
          # for node in nodes:  
    #     print(node.symbol, node.prob)
    
    # pick 2 smallest nodes
          right = nodes[0]
          left = nodes[1]
    
          left.code = 0
          right.code = 1
    
    # combine the 2 smallest nodes to create new node
          newNode = Node(left.prob+right.prob, left.symbol+right.symbol, left, right)
    
          nodes.remove(left)
          nodes.remove(right)
          nodes.append(newNode)
        
    huffman_encoding = Calculate_Codes(nodes[0])
    print("symbols with codes", huffman_encoding)
    Total_Gain(data, huffman_encoding)
    encoded_output = Output_Encoded(data,huffman_encoding)
    return encoded_output, nodes[0] 

encoding, tree = Huffman_Encoding(pixels)

file = open('compressed.txt','w')
file.write(encoding)
file.close()

内容的提问来源于stack exchange,提问作者hamid

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.17 02:50:51