You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python读取PGM图像遇Unicode解码错误,求解决方案

Fixing UnicodeDecodeError When Reading 16-bit PGM Files in Python

Hey, that UnicodeDecodeError you're hitting makes total sense—here's what's going on and how to fix it:

Your error comes from a core mismatch: 16-bit PGM files (usually the P5 binary format, not the ASCII P2 format your code assumes) store pixel data as raw binary bytes, not plain text. When you open the file with open(name) (text mode), Python tries to decode those binary bytes using your system's default charmap, which fails on non-text bytes like 0x90.

Here's a revised version of your code that handles both ASCII (P2) and binary (P5) PGM formats, including 16-bit variants:

import numpy as np
import matplotlib.pyplot as plt

def readpgm(name):
    with open(name, 'rb') as f:
        # Read and parse the PGM header (ASCII-encoded)
        header = []
        while len(header) < 4:
            line = f.readline().decode('ascii').strip()
            # Skip comment lines starting with #
            if line.startswith('#'):
                continue
            header.extend(line.split())
        
        # Validate format (supports both P2 and P5)
        assert header[0] in ('P2', 'P5'), "Unsupported PGM format"
        width, height = int(header[1]), int(header[2])
        max_val = int(header[3])
        
        # Determine data type: uint16 for 16-bit (max_val > 255), uint8 for 8-bit
        dtype = np.uint16 if max_val > 255 else np.uint8
        
        if header[0] == 'P2':
            # Handle ASCII PGM: read remaining text data
            pixel_data = []
            while len(pixel_data) < width * height:
                line = f.readline().decode('ascii').strip()
                if line.startswith('#'):
                    continue
                pixel_data.extend([int(val) for val in line.split()])
            img = np.array(pixel_data, dtype=dtype)
        else:
            # Handle binary PGM: read raw bytes
            # PGM specifies 16-bit data uses big-endian byte order
            img = np.frombuffer(f.read(), dtype='>' + str(dtype))
            # Ensure we read the correct number of pixels
            assert len(img) == width * height, "Mismatch in pixel count"
        
        # Reshape to 2D image array and return with max value
        return img.reshape((height, width)), max_val

# Usage example
img, max_val = readpgm('514516.pgm')
# Normalize display for 16-bit data (matplotlib works best with 0-1 or 0-255 ranges)
plt.imshow(img, cmap='gray', vmin=0, vmax=max_val)
plt.show()

Key Fixes & Explanations:

  • Binary File Mode: We open the file with 'rb' (read binary) to avoid automatic text decoding, which was causing your Unicode error.
  • Format Compatibility: The code checks for both P2 (ASCII) and P5 (binary) formats—critical because 16-bit PGMs almost always use the binary P5 format to save space.
  • 16-bit Data Handling: For binary P5 files, we use np.frombuffer to read raw bytes, specifying big-endian byte order (per PGM standards) to correctly interpret 16-bit pixel values.
  • Dynamic Data Type: The code automatically switches between uint8 (8-bit) and uint16 (16-bit) based on the max_val in the PGM header, so it works for both types of PGM files.

Why Your Original Code Failed:

Your code assumed the PGM was the ASCII-based P2 format, where all pixel values are stored as text. But 16-bit PGMs use the binary P5 format, where pixel data is stored as raw 2-byte values. Trying to read these binary bytes as text triggers the charmap decoding error you saw.

内容的提问来源于stack exchange,提问作者Nafis Islam

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 08:37:35