Python读取PGM图像遇Unicode解码错误,求解决方案
Hey, that UnicodeDecodeError you're hitting makes total sense—here's what's going on and how to fix it:
Your error comes from a core mismatch: 16-bit PGM files (usually the P5 binary format, not the ASCII P2 format your code assumes) store pixel data as raw binary bytes, not plain text. When you open the file with open(name) (text mode), Python tries to decode those binary bytes using your system's default charmap, which fails on non-text bytes like 0x90.
Here's a revised version of your code that handles both ASCII (P2) and binary (P5) PGM formats, including 16-bit variants:
import numpy as np import matplotlib.pyplot as plt def readpgm(name): with open(name, 'rb') as f: # Read and parse the PGM header (ASCII-encoded) header = [] while len(header) < 4: line = f.readline().decode('ascii').strip() # Skip comment lines starting with # if line.startswith('#'): continue header.extend(line.split()) # Validate format (supports both P2 and P5) assert header[0] in ('P2', 'P5'), "Unsupported PGM format" width, height = int(header[1]), int(header[2]) max_val = int(header[3]) # Determine data type: uint16 for 16-bit (max_val > 255), uint8 for 8-bit dtype = np.uint16 if max_val > 255 else np.uint8 if header[0] == 'P2': # Handle ASCII PGM: read remaining text data pixel_data = [] while len(pixel_data) < width * height: line = f.readline().decode('ascii').strip() if line.startswith('#'): continue pixel_data.extend([int(val) for val in line.split()]) img = np.array(pixel_data, dtype=dtype) else: # Handle binary PGM: read raw bytes # PGM specifies 16-bit data uses big-endian byte order img = np.frombuffer(f.read(), dtype='>' + str(dtype)) # Ensure we read the correct number of pixels assert len(img) == width * height, "Mismatch in pixel count" # Reshape to 2D image array and return with max value return img.reshape((height, width)), max_val # Usage example img, max_val = readpgm('514516.pgm') # Normalize display for 16-bit data (matplotlib works best with 0-1 or 0-255 ranges) plt.imshow(img, cmap='gray', vmin=0, vmax=max_val) plt.show()
Key Fixes & Explanations:
- Binary File Mode: We open the file with
'rb'(read binary) to avoid automatic text decoding, which was causing your Unicode error. - Format Compatibility: The code checks for both P2 (ASCII) and P5 (binary) formats—critical because 16-bit PGMs almost always use the binary P5 format to save space.
- 16-bit Data Handling: For binary P5 files, we use
np.frombufferto read raw bytes, specifying big-endian byte order (per PGM standards) to correctly interpret 16-bit pixel values. - Dynamic Data Type: The code automatically switches between
uint8(8-bit) anduint16(16-bit) based on themax_valin the PGM header, so it works for both types of PGM files.
Why Your Original Code Failed:
Your code assumed the PGM was the ASCII-based P2 format, where all pixel values are stored as text. But 16-bit PGMs use the binary P5 format, where pixel data is stored as raw 2-byte values. Trying to read these binary bytes as text triggers the charmap decoding error you saw.
内容的提问来源于stack exchange,提问作者Nafis Islam

