You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用fast_clustering.py聚类时出现长度不匹配错误求助

问题解决:聚类标签长度不匹配错误

错误原因

util.community_detection的返回结果不是每个样本对应的聚类标签,而是嵌套列表结构:每个子列表包含对应聚类的样本索引。你得到的17个元素代表17个聚类,每个元素是该聚类内所有样本的下标集合,直接赋值给DataFrame列会因长度不匹配触发报错。

解决方案

需要把聚类索引列表转换为与样本数匹配的标签数组:

  • 初始化一个长度等于样本数的数组,用-1标记未被聚类的样本(噪声点)
  • 遍历每个聚类,给对应索引位置分配聚类ID
  • 将转换后的标签数组赋值给DataFrame

修改后的完整代码

from sentence_transformers import SentenceTransformer, util
import pandas as pd
import time
import numpy as np
import torch

# 定义计算设备
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")

# 加载预训练模型
model = SentenceTransformer('paraphrase-MiniLM-L6-v2')

# 获取文本数据并生成嵌入
sentences = df['processed_activities'].tolist()
embeddings = model.encode(sentences)

# 转换为PyTorch张量并移至指定设备
embeddings = torch.from_numpy(embeddings).to(device)

print("Start clustering")
start_time = time.time()

# 聚类参数:min_community_size控制最小聚类规模,threshold控制相似度阈值
cluster_indices = util.community_detection(embeddings, min_community_size=25, threshold=0.75)

print("Clustering done after {:.2f} sec".format(time.time() - start_time))

# 将聚类索引转换为每个样本的标签
cluster_labels = np.full(len(sentences), -1)  # 初始化所有样本为未聚类状态
for cluster_id, indices in enumerate(cluster_indices):
    cluster_labels[indices] = cluster_id

# 给DataFrame添加聚类标签列
df['cluster'] = cluster_labels

# 打印所有有效聚类
num_clusters = df['cluster'].nunique() - 1  # 减去未聚类的-1
for i in range(num_clusters):
    print(f"Cluster {i}:")
    print(df.loc[df['cluster'] == i]['processed_activities'].values)

# 可选:查看未被聚类的样本
print("未聚类样本:")
print(df.loc[df['cluster'] == -1]['processed_activities'].values)

关键说明

  • 变量名改为cluster_indices,更准确反映其存储的是聚类索引集合
  • cluster_labels数组长度与样本数完全匹配,-1代表该样本未被分到任何符合最小规模要求的聚类中
  • 统计聚类数时排除了未聚类的标记,避免无效计数

内容的提问来源于stack exchange,提问作者Moe_blg

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.24 00:47:41