You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PDF解码:如何用PIL/Pillow还原经PNG Predictor算法处理的图像?

解决方案:PDF Predictor=15(PNG差分算法)解码问题

(a) Predictor=15对应的PNG过滤器算法及规范

Predictor=15对应PNG标准中的自适应行过滤机制,具体包含5种基础过滤器类型(由每行开头的1字节标识,取值0-4):

  • 0(None):无过滤,像素数据为原始值
  • 1(Sub):当前像素 = 差分数据 + 同当前行左侧像素的原始值
  • 2(Up):当前像素 = 差分数据 + 上一行同位置像素的原始值
  • 3(Average):当前像素 = 差分数据 + (左侧像素值 + 上一行同位置像素值) // 2
  • 4(Paeth):当前像素 = 差分数据 + Paeth预测器计算出的最优参考值(基于左侧、上方、左上三个像素)

编码时会为每行选择压缩效率最高的过滤器,解码时必须读取每行开头的过滤器标识,再逆向计算出原始像素数据。该逻辑定义在PNG标准(ISO/IEC 15948)的第9章。

(b) 实现还原的方案

1. 使用第三方库:pypng

pypng库原生支持处理带PNG过滤器的扫描线数据,无需手动实现过滤逻辑:

pip install pypng

修改你的代码,替换图像生成部分:

import png

@property
def image(self):
    im = self._image
    if im is None:
        decoded_bs = self.decoded_payload
        decode_params = self.context_dict.get(b'DecodeParms', {})
        color_space = self.context_dict[b'ColorSpace']
        bits_per_component = decode_params.get(b'BitsPerComponent') or {b'DeviceRGB':8, b'DeviceGray':8}[color_space]
        colors = decode_params.get(b'Colors') or {b'DeviceRGB':3, b'DeviceGray':1}[color_space]
        width = self.context_dict[b'Width']
        height = self.context_dict[b'Height']
        
        # 计算每行像素字节数(不含开头的过滤器字节)
        row_pixel_bytes = ((width * bits_per_component * colors) + 7) // 8
        # 拆分扫描线:每行1字节过滤器 + row_pixel_bytes字节数据
        scanlines = []
        offset = 0
        for _ in range(height):
            filter_byte = decoded_bs[offset]
            row_data = decoded_bs[offset+1 : offset+1+row_pixel_bytes]
            scanlines.append(row_data)
            offset += 1 + row_pixel_bytes
        
        # 用pypng解码带过滤器的扫描线,转换为PIL图像
        reader = png.Reader(
            width=width,
            height=height,
            bitdepth=bits_per_component,
            channels=colors,
            interlace=0,
            filter='all',
            scanlines=scanlines
        )
        png_data = reader.asDirect()
        pixels = png_data[2]
        mode = 'L' if colors ==1 else 'RGB'
        im = Image.frombytes(mode, (width, height), b''.join(pixels))
        
        self._image = im
    return im

2. 纯Python手动实现过滤器

如果不想依赖第三方库,可按PNG过滤器规则逐行逆向计算:

@property
def image(self):
    im = self._image
    if im is None:
        decoded_bs = self.decoded_payload
        decode_params = self.context_dict.get(b'DecodeParms', {})
        color_space = self.context_dict[b'ColorSpace']
        bits_per_component = decode_params.get(b'BitsPerComponent') or {b'DeviceRGB':8, b'DeviceGray':8}[color_space]
        colors = decode_params.get(b'Colors') or {b'DeviceRGB':3, b'DeviceGray':1}[color_space]
        width = self.context_dict[b'Width']
        height = self.context_dict[b'Height']
        
        # 计算每个像素的字节数、每行像素总字节数
        bytes_per_pixel = (bits_per_component * colors +7)//8
        row_pixel_bytes = width * bytes_per_pixel
        output = bytearray()
        prev_row = bytearray(row_pixel_bytes)  # 存储上一行的原始像素数据
        
        offset =0
        for _ in range(height):
            filter_type = decoded_bs[offset]
            row_data = decoded_bs[offset+1 : offset+1+row_pixel_bytes]
            offset +=1 + row_pixel_bytes
            
            current_row = bytearray(row_pixel_bytes)
            if filter_type ==0:
                # None过滤器:直接复制
                current_row[:] = row_data
            elif filter_type ==1:
                # Sub过滤器:current = data + left
                for i in range(row_pixel_bytes):
                    left = current_row[i - bytes_per_pixel] if i >= bytes_per_pixel else 0
                    current_row[i] = (row_data[i] + left) & 0xFF
            elif filter_type ==2:
                # Up过滤器:current = data + up
                for i in range(row_pixel_bytes):
                    up = prev_row[i]
                    current_row[i] = (row_data[i] + up) &0xFF
            elif filter_type ==3:
                # Average过滤器:current = data + (left + up)//2
                for i in range(row_pixel_bytes):
                    left = current_row[i - bytes_per_pixel] if i >= bytes_per_pixel else0
                    up = prev_row[i]
                    current_row[i] = (row_data[i] + (left + up)//2) &0xFF
            elif filter_type ==4:
                # Paeth过滤器:current = data + Paeth预测值
                def paeth(a,b,c):
                    p = a + b -c
                    pa = abs(p -a)
                    pb = abs(p -b)
                    pc = abs(p -c)
                    if pa <= pb and pa <= pc:
                        return a
                    elif pb <= pc:
                        return b
                    else:
                        return c
                for i in range(row_pixel_bytes):
                    left = current_row[i - bytes_per_pixel] if i >= bytes_per_pixel else0
                    up = prev_row[i]
                    up_left = prev_row[i - bytes_per_pixel] if i >= bytes_per_pixel else0
                    pred = paeth(left, up, up_left)
                    current_row[i] = (row_data[i] + pred) &0xFF
            
            output.extend(current_row)
            prev_row = current_row
        
        # 生成PIL图像
        PIL_mode = {
            (b'DeviceGray', 1,1,0): 'L',
            (b'DeviceGray',8,1,0): 'L',
            (b'DeviceRGB',8,3,0): 'RGB'
        }[(color_space, bits_per_component, colors, decode_params.get(b'ColorTransform',0))]
        im = Image.frombytes(PIL_mode, (width, height), output)
        self._image = im
    return im

问题根源说明

你当前的代码直接将带过滤器标识的差分数据传给Image.frombytes,导致:

  1. 每行开头的1字节过滤器标识被当作像素数据,使每行实际长度多1字节,整体图像偏移错位
  2. 差分数据未被还原为原始像素值,导致颜色/亮度异常

内容的提问来源于stack exchange,提问作者Cameron Simpson

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.10 12:44:51