You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

KMeans图像聚类循环失效:标签与图片数量不匹配问题求助

图像聚类KMeans标签匹配问题修复

问题场景

使用GitHub仓库study2A的代码进行图像聚类时,KMeans聚类环节无法为每张图片分配标签——生成的标签数量为9,但图片数量为200,二者长度不匹配,无法完成标签赋值并生成带标签的txt文件。

原代码

import joblib, os
import numpy as np
import pandas as pd
from sklearn.cluster import KMeans
import ypoften as of

cwd = os.path.join('/content/drive/MyDrive/study2A/for lab teaching',"")
# please change this to the folder that has all the images

cvmodels = ["vgg16 fc1"]
## cvmodels = ["vgg16 fc1","vgg16 places fc1","vggface fc6"] choose the model you used
clmethod = "KMeans"

for cvmodel in cvmodels:

    features_savepath = os.path.join(cwd,'img exfeature1','features PCA',cvmodel+'.dat')
    features_array = joblib.load(features_savepath)
    X = pd.DataFrame(features_array)
    print(X)

    nd = 200
    X = X.iloc[:,0:nd]

    savefolder = cvmodel + ' ' + clmethod + ' PCA' + str(nd)

    # Read the 0321_imageselected_A.txt file
    imgnamefile = os.path.join(cwd, "0321_imageselected_A.txt")
    df = pd.read_csv(imgnamefile, sep='\t', header=0)

    for K in [2]:
        print('number of cluster', K)
        cl = KMeans(K, random_state=0)
        cl.fit(X)
        labels = cl.labels_

        imgnamefile = os.path.join(cwd,"0321_imageselected_A.txt")
        df = pd.read_csv(imgnamefile, sep ='\t', header = 0)
        print(df)

        df['label'] = labels
        filepath = os.path.join(cwd,'img cluster',savefolder,str(K),'label.txt')
        of.create_path(filepath)
        df.to_csv(filepath, index = None, header = None, sep = '\t')

    print("DONE"*20)

报错信息

ValueError: Length of values (9) does not match length of index (200)

原因分析

报错的核心是加载的特征矩阵features_array仅包含9个样本的特征,而图片列表0321_imageselected_A.txt中有200张图片。KMeans聚类后生成的labels长度等于特征矩阵的行数(9),但图片数据框df的行数是200,赋值时必然出现长度不匹配。

修复步骤

  1. 验证特征矩阵样本数:在加载特征文件后,打印特征矩阵的形状,确认样本数是否与图片数量一致。
  2. 排查特征提取环节:如果特征矩阵样本数确实为9,说明特征提取阶段未正确为200张图片生成特征,需要回到特征提取代码中检查逻辑,确保生成的.dat文件包含200个样本的特征数据。
  3. 优化代码冗余:原代码中重复读取图片列表文件,可简化以提升效率。

修复后代码

import joblib, os
import numpy as np
import pandas as pd
from sklearn.cluster import KMeans
import ypoften as of

cwd = os.path.join('/content/drive/MyDrive/study2A/for lab teaching',"")

cvmodels = ["vgg16 fc1"]
clmethod = "KMeans"

for cvmodel in cvmodels:
    features_savepath = os.path.join(cwd,'img exfeature1','features PCA',cvmodel+'.dat')
    features_array = joblib.load(features_savepath)
    
    # 关键:打印特征矩阵形状,验证样本数是否为200
    print(f"特征矩阵形状(样本数,特征数):{features_array.shape}")
    if features_array.shape[0] != 200:
        print("警告:特征样本数与图片数量不匹配,请检查特征提取环节!")
        continue

    X = pd.DataFrame(features_array)
    nd = 200
    X = X.iloc[:,0:nd]

    savefolder = cvmodel + ' ' + clmethod + ' PCA' + str(nd)

    # 仅读取一次图片列表
    imgnamefile = os.path.join(cwd, "0321_imageselected_A.txt")
    df = pd.read_csv(imgnamefile, sep='\t', header=0)

    for K in [2]:
        print('聚类数量:', K)
        cl = KMeans(K, random_state=0)
        cl.fit(X)
        labels = cl.labels_

        # 确认标签长度与图片数量一致
        print(f"标签长度: {len(labels)}, 图片数量: {len(df)}")
        df['label'] = labels

        filepath = os.path.join(cwd,'img cluster',savefolder,str(K),'label.txt')
        of.create_path(filepath)
        df.to_csv(filepath, index=None, header=None, sep='\t')

    print("DONE"*20)

注意事项

  • 如果打印特征矩阵形状后发现样本数不是200,必须优先修正特征提取流程,确保每个图片对应一行特征数据。
  • 若特征矩阵样本数正确但仍报错,可检查图片列表文件是否存在重复行或读取错误。

内容的提问来源于stack exchange,提问作者user23713357

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.27 16:13:14