使用fast_clustering.py聚类时出现长度不匹配错误求助
问题解决:聚类标签长度不匹配错误
错误原因
util.community_detection的返回结果不是每个样本对应的聚类标签,而是嵌套列表结构:每个子列表包含对应聚类的样本索引。你得到的17个元素代表17个聚类,每个元素是该聚类内所有样本的下标集合,直接赋值给DataFrame列会因长度不匹配触发报错。
解决方案
需要把聚类索引列表转换为与样本数匹配的标签数组:
- 初始化一个长度等于样本数的数组,用
-1标记未被聚类的样本(噪声点) - 遍历每个聚类,给对应索引位置分配聚类ID
- 将转换后的标签数组赋值给DataFrame
修改后的完整代码
from sentence_transformers import SentenceTransformer, util import pandas as pd import time import numpy as np import torch # 定义计算设备 device = torch.device("cuda" if torch.cuda.is_available() else "cpu") # 加载预训练模型 model = SentenceTransformer('paraphrase-MiniLM-L6-v2') # 获取文本数据并生成嵌入 sentences = df['processed_activities'].tolist() embeddings = model.encode(sentences) # 转换为PyTorch张量并移至指定设备 embeddings = torch.from_numpy(embeddings).to(device) print("Start clustering") start_time = time.time() # 聚类参数:min_community_size控制最小聚类规模,threshold控制相似度阈值 cluster_indices = util.community_detection(embeddings, min_community_size=25, threshold=0.75) print("Clustering done after {:.2f} sec".format(time.time() - start_time)) # 将聚类索引转换为每个样本的标签 cluster_labels = np.full(len(sentences), -1) # 初始化所有样本为未聚类状态 for cluster_id, indices in enumerate(cluster_indices): cluster_labels[indices] = cluster_id # 给DataFrame添加聚类标签列 df['cluster'] = cluster_labels # 打印所有有效聚类 num_clusters = df['cluster'].nunique() - 1 # 减去未聚类的-1 for i in range(num_clusters): print(f"Cluster {i}:") print(df.loc[df['cluster'] == i]['processed_activities'].values) # 可选:查看未被聚类的样本 print("未聚类样本:") print(df.loc[df['cluster'] == -1]['processed_activities'].values)
关键说明
- 变量名改为
cluster_indices,更准确反映其存储的是聚类索引集合 cluster_labels数组长度与样本数完全匹配,-1代表该样本未被分到任何符合最小规模要求的聚类中- 统计聚类数时排除了未聚类的标记,避免无效计数
内容的提问来源于stack exchange,提问作者Moe_blg
相关产品推荐
相关产品推荐

