为什么执行PCA降维操作后,生成的图像文件大小反而比原图更大?
PCA处理后图像体积反而变大问题排查
问题描述
我正在构建针对美国鹿类物种的图像分类模型,目前计划对数据集图像执行PCA操作以降低内存占用、减少后续模型运行耗时。
我了解主成分分析(PCA)的作用是在尽量保留数据集方差的前提下降低维度,因此当我发现经过自定义Deer_PCA函数处理后的所有PCA压缩图像大小均大于原图时十分困惑:当设置n_components = 150时,原图大小为128 KB,处理后的新图像大小反而达到了293 KB,有人知道出现这个问题的原因是什么吗?
参考图像
- 输入原图:

- PCA处理后输出图像:

相关代码
#import some packages import cv2 import os,sys from PIL import Image import pandas as pd from scipy.stats import stats from sklearn.decomposition import PCA import matplotlib.pyplot as plt import matplotlib.image as mpimg #Let's write a function to perform PCA on all the images in the folder and output it to a new folder #inpath = folder containing the image - string value #outpath = which folder do I want the new compressed image saved to. - string value #n_comp = number of components - int value def Deer_PCA(inpath, outpath,n_comp): for image_path in os.listdir(inpath): # create the full input path and read the file input_path = os.path.join(inpath, image_path) print(input_path) w_deer = cv2.cvtColor(cv2.imread(input_path), cv2.COLOR_BGR2RGB) #split image blue_2,green_2,red_2 = cv2.split(w_deer) #scale channels w_blue = blue_2/255 w_green = green_2/255 w_red = red_2/255 #PCA on each channel pca_b2 = PCA(n_components=n_comp) pca_b2.fit(w_blue) trans_pca_b2 = pca_b2.transform(w_blue) pca_g2 = PCA(n_components=n_comp) pca_g2.fit(w_green) trans_pca_g2 = pca_g2.transform(w_green) pca_r2 = PCA(n_components=n_comp) pca_r2.fit(w_red) trans_pca_r2 = pca_r2.transform(w_red) #merge channels after PCA b_arr2 = pca_b2.inverse_transform(trans_pca_b2) g_arr2 = pca_g2.inverse_transform(trans_pca_g2) r_arr2 = pca_r2.inverse_transform(trans_pca_r2) img_reduced2 = (cv2.merge((b_arr2, g_arr2, r_arr2))) print("Merge Successful") # create the full output path fullpath = os.path.join(outpath, 'PCA_'+image_path) cv2.imwrite(fullpath, img_reduced2*255) print("Successfully saved\n") #Check the image sizes original_image_path = '/Users/matthew_macwan/Downloads/CIS/I_Class_Deer/mule_deer_doe/mule deer doe_1.jpeg' PCA_compressed_image_path = '/Users/matthew_macwan/Downloads/CIS/I_Class_Deer/mule_deer_doe/PCA_mule deer doe_1.jpeg' print('Original Image:',sys.getsizeof(original_image_path)) print('PCA Image:',sys.getsizeof(PCA_compressed_image_path))
问题原因
- 体积统计逻辑错误:你使用的
sys.getsizeof()统计的是文件路径字符串的内存占用,不是图片文件的磁盘体积,正确统计文件大小应该用os.path.getsize(文件路径)。 - 对PCA压缩的作用场景理解错误:PCA的压缩效果体现在降维后得到的
trans_pca_b2、trans_pca_g2、trans_pca_r2这些特征矩阵上,你把特征矩阵逆变换重建回像素空间后,再存为JPEG格式时,重建后的图像丢失了原图的高频冗余特性,JPEG算法的压缩效率会大幅下降,最终输出文件体积反而比原图更大。
优化建议
- 如果做PCA是为了给后续分类模型输入特征,不需要保存逆变换后的图像,直接存储降维后的PCA特征矩阵即可,这才是真正降低内存占用、减少后续计算量的用法。
- 如果必须保存重建后的可视化图像,可以调整JPEG保存的质量参数,例如:
cv2.imwrite(fullpath, img_reduced2*255, [int(cv2.IMWRITE_JPEG_QUALITY), 85])
通过调低质量参数控制输出文件体积。
内容的提问来源于stack exchange,提问作者Matthew
相关产品推荐
相关产品推荐

