You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python Autoencoder图像压缩输出图像混乱问题求助

图像Autoencoder重构结果混乱的问题排查与解决

问题描述

尝试用Python构建Autoencoder实现图像压缩,但重构图像始终呈现像素混合错乱的状态,无法还原原图特征。

代码实现

import os
import cv2
import numpy as np
import tensorflow as tf
from tensorflow.keras import Model, layers, losses
import matplotlib.pyplot as plt

path1 = 'C:\\Users\\klaud\\Desktop\\images\\'
all_images = []
subjects = os.listdir(path1)
numberOfSubject = len(subjects)
print('Number of Subjects: ', numberOfSubject)
for number1 in range(0, numberOfSubject):  # numberOfSubject
    path2 = (path1 + subjects[number1] + '/')
    sequences = os.listdir(path2)
    numberOfsequences = len(sequences)
    for number2 in range(0, numberOfsequences):
        path3 = path2 + sequences[number2]
        img = cv2.imread(path3, 0)
        img = img.reshape(512, 512, 1)
        all_images.append(img)

x_train = np.array([all_images[0], all_images[1]])
x_test = np.array(all_images[2:])

print("X TRAIN \n")
print(x_train)
print("X TEST \n")
print(x_test)

x_train = x_train.astype('float32') / 255.
x_test = x_test.astype('float32') / 255.

print (x_train.shape)
print (x_test.shape)

latent_dim = 4

class Autoencoder(Model):
    def __init__(self, latent_dim):
        super(Autoencoder, self).__init__()
        self.latent_dim = latent_dim
        self.encoder = tf.keras.Sequential([
            layers.Flatten(),
            layers.Dense(latent_dim, activation='relu'),
        ])
        self.decoder = tf.keras.Sequential([
            layers.Dense(262144, activation='sigmoid'),
            layers.Reshape((512, 512))
        ])

    def call(self, x):
        encoded = self.encoder(x)
        decoded = self.decoder(encoded)
        return decoded

autoencoder = Autoencoder(latent_dim)

autoencoder.compile(optimizer='adam', loss=losses.MeanSquaredError())

autoencoder.fit(x_train, x_train,
                epochs=10,
                shuffle=True,
                validation_data=(x_test, x_test))

encoded_imgs = autoencoder.encoder(x_test).numpy()
decoded_imgs = autoencoder.decoder(encoded_imgs).numpy()

n = 6
plt.figure(figsize=(20, 6))
for i in range(n):
    # display original
    ax = plt.subplot(2, n, i + 1)
    plt.imshow(x_test[i])
    plt.title("original")
    plt.gray()
    ax.get_xaxis().set_visible(False)
    ax.get_yaxis().set_visible(False)

    # display reconstruction
    ax = plt.subplot(2, n, i + 1 + n)
    plt.imshow(decoded_imgs[i])
    plt.title("reconstructed")
    plt.gray()
    ax.get_xaxis().set_visible(False)
    ax.get_yaxis().set_visible(False)
plt.show()

核心问题分析与解决建议

1. 训练数据量严重不足

  • 仅用2张图做训练集,模型无法学习到通用的图像特征分布。Autoencoder需要足够多样本捕捉图像模式,建议使用几十到上百张同类图像训练。
  • 调整训练集划分:x_train = np.array(all_images[:int(len(all_images)*0.8)]),x_test = np.array(all_images[int(len(all_images)*0.8):]),按比例分配训练/测试数据。

2. 潜在维度设置过小

  • latent_dim=4对512x512图像的压缩比过高(262144→4,压缩65536倍),模型无法保留足够重构信息。
  • 先尝试将latent_dim提升到256或512,后续再逐步调整找到合适的压缩平衡点。

3. 模型结构过于简单

  • 单层全连接层无法提取复杂图像特征,改用卷积Autoencoder能更好捕捉空间局部特征:
class Autoencoder(Model):
    def __init__(self, latent_dim):
        super(Autoencoder, self).__init__()
        self.latent_dim = latent_dim
        # 编码器:卷积+池化
        self.encoder = tf.keras.Sequential([
            layers.Input(shape=(512, 512, 1)),
            layers.Conv2D(32, (3, 3), activation='relu', padding='same', strides=2),
            layers.Conv2D(64, (3, 3), activation='relu', padding='same', strides=2),
            layers.Flatten(),
            layers.Dense(latent_dim, activation='relu')
        ])
        # 解码器:反卷积+上采样
        self.decoder = tf.keras.Sequential([
            layers.Dense(128*128*64, activation='relu'),
            layers.Reshape((128, 128, 64)),
            layers.Conv2DTranspose(64, (3, 3), activation='relu', padding='same', strides=2),
            layers.Conv2DTranspose(32, (3, 3), activation='relu', padding='same', strides=2),
            layers.Conv2D(1, (3, 3), activation='sigmoid', padding='same')
        ])

    def call(self, x):
        encoded = self.encoder(x)
        decoded = self.decoder(encoded)
        return decoded

4. 绘图维度不匹配

  • 原图像为(512,512,1)单通道格式,plt.imshow默认更适配(H,W)格式,显式去掉通道维度可避免像素排列异常:
    • 显示原图:plt.imshow(x_test[i].squeeze())
    • 显示重构图:plt.imshow(decoded_imgs[i].squeeze())

5. 训练轮次不足

  • 仅训练10轮无法让模型充分收敛,建议将epochs提升到50-100轮,同时加入早停机制防止过拟合:
from tensorflow.keras.callbacks import EarlyStopping
es = EarlyStopping(monitor='val_loss', patience=5, restore_best_weights=True)
autoencoder.fit(x_train, x_train,
                epochs=100,
                shuffle=True,
                validation_data=(x_test, x_test),
                callbacks=[es])

内容的提问来源于stack exchange,提问作者clautsick

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.02 10:10:31