如何计算图像数据集不同类别间的相关性?
问题描述
我尝试计算不同类别间的相关性。我的数据集包含多个以类别命名的文件夹,每个文件夹下有若干图像,已使用以下代码读取图像路径并完成标注:
# Specify root path root_path = './Datasets_1' name_class = os.listdir(root_path) dataset_Path = list(glob.glob(root_path+'/**/*.*')) labels = list(map(lambda x : os.path.split(os.path.split(x)[0])[1],dataset_Path)) dataset_Path = pd.Series(dataset_Path,name = 'FilePath').astype(str) labels = pd.Series(labels,name='Label') data = pd.concat([dataset_Path,labels],axis =1) data = data.sample(frac=1).reset_index(drop= True)
请问如何计算这些类别之间的相关性?
解决方案
要计算图像类别间的相关性,核心是先把图像转换成数值特征,再基于类别特征的统计量计算类别间的关联度,具体步骤如下:
1. 提取图像数值特征
图像本身是像素矩阵,无法直接计算相关性,需要转换成可量化的特征向量,推荐两种常用方法:
方法1:预训练CNN提取语义特征
用深度学习模型提取高层语义特征,适合需要类别语义关联的场景:
import torch import torchvision.models as models import torchvision.transforms as transforms from PIL import Image # 加载预训练ResNet50,去掉最后全连接分类层 model = models.resnet50(pretrained=True) feature_extractor = torch.nn.Sequential(*list(model.children())[:-1]) feature_extractor.eval() # 定义图像预处理流程(与模型训练时一致) transform = transforms.Compose([ transforms.Resize((224, 224)), transforms.ToTensor(), transforms.Normalize(mean=[0.485, 0.456, 0.406], std=[0.229, 0.224, 0.225]) ]) # 单张图像特征提取函数 def extract_features(img_path): img = Image.open(img_path).convert('RGB') img_tensor = transform(img).unsqueeze(0) with torch.no_grad(): # 关闭梯度计算,节省资源 features = feature_extractor(img_tensor) return features.flatten().numpy() # 展平成一维特征向量 # 为所有图像提取特征 data['Features'] = data['FilePath'].apply(extract_features)
方法2:传统视觉特征(如颜色直方图)
如果不需要语义特征,可使用颜色、纹理等传统特征:
import cv2 import numpy as np # 提取HSV颜色直方图特征 def extract_color_hist(img_path): img = cv2.imread(img_path) hsv = cv2.cvtColor(img, cv2.COLOR_BGR2HSV) # 计算三维直方图,分箱数为8x8x8 hist = cv2.calcHist([hsv], [0,1,2], None, [8,8,8], [0,256, 0,256, 0,256]) cv2.normalize(hist, hist) # 归一化 return hist.flatten() data['Features'] = data['FilePath'].apply(extract_color_hist)
2. 计算类别特征均值
每个类别下的图像特征取均值,得到该类别的代表特征向量:
import numpy as np # 按类别分组,堆叠所有图像特征后计算均值 class_features = data.groupby('Label')['Features'].apply(lambda x: np.mean(np.stack(x), axis=0)) # 转换成DataFrame,方便后续计算 class_features_df = pd.DataFrame(class_features.tolist(), index=class_features.index)
3. 计算类别间相关性矩阵
基于类别特征均值矩阵,计算皮尔逊相关系数(最常用的相关性指标):
# 计算皮尔逊相关系数矩阵 correlation_matrix = class_features_df.T.corr() # 可视化相关性矩阵(可选) import seaborn as sns import matplotlib.pyplot as plt plt.figure(figsize=(10,8)) sns.heatmap(correlation_matrix, annot=True, cmap='coolwarm', fmt='.2f') plt.title('Category Correlation Matrix') plt.show()
补充说明
- 皮尔逊相关系数取值范围为[-1,1]:越接近1表示两类特征越相似,越接近-1表示差异越大
- 若更关注特征向量的方向相似度,可替换为余弦相似度:
from sklearn.metrics.pairwise import cosine_similarity cosine_matrix = cosine_similarity(class_features_df) - 预训练模型可根据需求替换,比如EfficientNet、ViT等,语义特征的表达能力会有所不同
内容的提问来源于stack exchange,提问作者user1967441
相关产品推荐
相关产品推荐

