You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

求助:基于Principal Component Analysis(PCA)的5类图像分类代码实现

使用PCA实现RGB图像分类的代码方案

依赖库安装

先安装所需的Python库:

pip install numpy scikit-learn pillow matplotlib

完整代码实现

1. 数据加载与预处理

import os
import numpy as np
from PIL import Image
from sklearn.decomposition import PCA
from sklearn.model_selection import train_test_split
from sklearn.svm import SVC
from sklearn.metrics import classification_report, confusion_matrix
from sklearn.preprocessing import StandardScaler
import matplotlib.pyplot as plt

def load_images(data_dir, img_size=(224, 224)):
    """加载图像数据,返回特征矩阵、标签和类别名称"""
    X = []
    y = []
    class_names = sorted(os.listdir(data_dir))
    
    for class_idx, class_name in enumerate(class_names):
        class_dir = os.path.join(data_dir, class_name)
        if not os.path.isdir(class_dir):
            continue
        
        for img_name in os.listdir(class_dir):
            img_path = os.path.join(class_dir, img_name)
            try:
                # 读取RGB图像并统一尺寸
                img = Image.open(img_path).convert('RGB').resize(img_size)
                # 展平成一维特征向量
                img_flat = np.array(img).flatten()
                X.append(img_flat)
                y.append(class_idx)
            except Exception as e:
                print(f"跳过损坏图像 {img_path}: {e}")
    
    return np.array(X), np.array(y), class_names

2. PCA降维与分类主逻辑

if __name__ == "__main__":
    # 配置参数(替换成你的实际路径和需求)
    DATA_DIR = "path/to/your/dataset"  # 数据集根目录,子文件夹为类别
    IMG_SIZE = (224, 224)              # 统一图像尺寸
    PCA_VARIANCE_RATIO = 0.95          # PCA保留的方差比例

    # 加载数据集
    X, y, class_names = load_images(DATA_DIR, IMG_SIZE)
    print(f"已加载 {X.shape[0]} 张图像,原始特征维度:{X.shape[1]}")

    # 划分训练集/测试集(按类别分层抽样)
    X_train, X_test, y_train, y_test = train_test_split(
        X, y, test_size=0.2, random_state=42, stratify=y
    )

    # 数据标准化(PCA对尺度敏感,必须执行)
    scaler = StandardScaler()
    X_train_scaled = scaler.fit_transform(X_train)
    X_test_scaled = scaler.transform(X_test)

    # 执行PCA降维
    pca = PCA(n_components=PCA_VARIANCE_RATIO)
    X_train_pca = pca.fit_transform(X_train_scaled)
    X_test_pca = pca.transform(X_test_scaled)
    print(f"PCA降维后训练集维度:{X_train_pca.shape}")
    print(f"保留主成分数量:{pca.n_components_},累计解释方差:{sum(pca.explained_variance_ratio_):.2f}")

    # 训练SVM分类器(可替换为其他分类器)
    classifier = SVC(kernel='rbf', C=1.0, random_state=42)
    classifier.fit(X_train_pca, y_train)

    # 模型评估
    y_pred = classifier.predict(X_test_pca)
    print("\n===== 分类报告 =====")
    print(classification_report(y_test, y_pred, target_names=class_names))
    
    print("\n===== 混淆矩阵 =====")
    print(confusion_matrix(y_test, y_pred))

    # 可选:可视化前5个主成分对应的特征图
    plt.figure(figsize=(12, 4))
    for i in range(min(5, pca.n_components_)):
        # 将主成分重塑为图像尺寸并标准化到0-255
        component = pca.components_[i].reshape(IMG_SIZE[0], IMG_SIZE[1], 3)
        component = (component - component.min()) / (component.max() - component.min()) * 255
        component = component.astype(np.uint8)
        
        plt.subplot(1, 5, i+1)
        plt.imshow(component)
        plt.title(f"Component {i+1}")
        plt.axis('off')
    plt.tight_layout()
    plt.show()

关键注意事项

  • 数据集结构:确保你的数据集按类别分文件夹存放,比如 dataset/class_a、dataset/class_b,load_images 会自动识别子文件夹作为类别。
  • 图像尺寸调整:如果原始图像尺寸差异大,必须统一resize;减小尺寸(比如(128,128))可以大幅降低特征维度,加快运算速度,但可能损失细节。
  • PCA参数调整:
    • 设为小数(如0.95):自动选择主成分数量以保留对应比例的方差,适合需要平衡维度和信息的场景。
    • 设为整数:直接指定主成分数量,适合固定输出维度的需求。
  • 分类器替换:除了SVM,你也可以尝试 LogisticRegression、RandomForestClassifier 等,根据实际分类效果选择。
  • 标准化的必要性:RGB像素值范围是0-255,不同通道的方差差异会导致PCA偏向方差大的通道,标准化后每个特征均值为0、方差为1,保证PCA的公平性。

内容的提问来源于stack exchange,提问作者rayan matlob

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.08 16:15:43