You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PGM灰度图膨胀算法实现异常及行列赋值问题求助

解决PGM图像处理中的膨胀异常与赋值问题

嘿,我帮你排查出了代码里的核心问题,都是源于PGM文件读取时的行列顺序搞反了,咱们一步步来解决:

问题1:膨胀操作生成多个对象副本

你的pgm_read函数在解析PGM尺寸时犯了个常见错误:PGM P2格式的第三行是宽度(列数) 高度(行数),但numpy的reshape是按**(行数, 列数)**来构建数组的。你当前的代码把宽度当成行数、高度当成列数,导致整个图像被转置了——后续的膨胀操作相当于在转置后的图像上执行,自然会出现不符合预期的对象副本。

修复pgm_read函数

修改reshape的参数顺序,把高度作为行数、宽度作为列数:

def pgm_read(filename):
    """Read PGM file to a array"""
    try:
        with open(filename, 'r') as fp:
            lines = fp.readlines()
            header_info = lines[2].split()
            width = int(header_info[0])
            height = int(header_info[1])
            # 按(高度, 宽度)的顺序reshape,对应numpy的行x列结构
            pixel_data = np.array([line.strip('\n') for line in lines[4:]], dtype=np.int32)
            return pixel_data.reshape(height, width)
    except OSError:
        print("An exception occurred")

修复后,你的膨胀逻辑就能正常工作了:每个原始前景像素会把2x2核覆盖的邻域都设为前景,最终实现对象的放大,而不是产生奇怪的副本。

问题2:左半部分赋值异常

同样是图像被转置的锅:你代码里的rows其实是原图像的列数,columns是原图像的行数。当你执行for y in range(columns // 2)时,实际修改的是转置后图像的前半列,对应原图像的前半行,完全不是你预期的左半部分(原图像的前半列)。

修复行列顺序后,再执行左半部分赋值代码:

rows, columns = arr.shape
# 现在的rows是原图像行数,columns是原图像列数,赋值逻辑正确
for x in range(rows):
    for y in range(columns // 2):
        arr[x][y] = 80

这时候就能正确修改原图像的左半部分(所有行的前半列)了。

额外的效率优化建议

  • 你的pgm_write逐像素写入的方式效率很低,可以改成逐行拼接写入:
def pgm_write(img, dest, header):
    """Write numpy array to PGM file"""
    try:
        width, height = img.shape[1], img.shape[0]
        header = f"P2\n# test\n{width} {height}\n80\n"
        with open(dest, "w") as f:
            f.write(header)
            # 逐行写入像素,用空格分隔
            for row in img:
                f.write(' '.join(map(str, row)) + '\n')
    except OSError:
        print("Writing exception occurred")
  • 膨胀操作可以用numpy向量化替代嵌套循环,速度会快很多:
def dilation(img):
    rows, cols = img.shape
    dilated_img = np.zeros((rows, cols), dtype=np.int32)
    kernel = np.ones((2, 2), dtype=np.bool_)
    # 获取所有前景像素的坐标
    y_coords, x_coords = np.where(img == 1)
    # 遍历核的偏移量
    for i in range(kernel.shape[0]):
        for j in range(kernel.shape[1]):
            if kernel[i][j]:
                new_y = y_coords + i
                new_x = x_coords + j
                # 过滤超出图像范围的坐标
                valid_mask = (new_y < rows) & (new_x < cols)
                dilated_img[new_y[valid_mask], new_x[valid_mask]] = 1
    # 替换为高亮值用于调试
    dilated_img[dilated_img == 1] = 80
    return dilated_img

内容的提问来源于stack exchange,提问作者Melvyn Tie

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.06 15:39:10