基于t-SNE与K-Means的无监督学习报错排查及降维疑问
问题背景
我正在数据集上开展无监督学习,目标是提取特征、找出聚类分组及各簇核心特征(centroid)。计划先用t-SNE降低数据维度以便绘制散点图,再通过K-Means获取各簇中心权重,最终找到包含最多Bad状态样本、最少Good状态样本的簇中心。
示例代码如下:
# Set a seed for reproducibility np.random.seed(42) # Generate dummy data with random values num_rows = 1000 # Create a DataFrame with random values and specific column names dummy_data = pd.DataFrame({ 'Name': [np.random.choice(['Alice', 'Bob', 'Charlie', 'David', 'Eva', 'Frank', 'Grace', 'Henry', 'Isabella', 'Jack', 'Kate', 'Liam', 'Mia', 'Noah', 'Olivia', 'Peter', 'Quinn', 'Rachel', 'Sam', 'Taylor'] ) for _ in range(num_rows)], 'Condition': np.random.choice(['Good', 'Bad'], size=num_rows), 'Latency_Wifi': np.random.normal(loc=1, scale=0.2, size=num_rows), # 'Good' condition has lower latency 'Loss_Wifi': np.random.normal(loc=0.05, scale=0.02, size=num_rows), # 'Good' condition has lower loss 'Latency_Gaming': np.random.normal(loc=1, scale=0.2, size=num_rows), 'Loss_Gaming': np.random.normal(loc=0.05, scale=0.02, size=num_rows), 'Latency_Video': np.random.normal(loc=1, scale=0.2, size=num_rows), 'Loss_Video': np.random.normal(loc=0.05, scale=0.02, size=num_rows), 'Latency_WFH': np.random.normal(loc=1, scale=0.2, size=num_rows), 'Loss_WFH': np.random.normal(loc=0.05, scale=0.02, size=num_rows), }) features = dummy_data.drop(['Name', 'Condition'], axis=1) # Standardize the data to have zero mean and unit variance scaler = StandardScaler() data_scaled = scaler.fit_transform(features) # kpca = KernelPCA(n_components=10, kernel='rbf', gamma=0.1) # data_kpca = kpca.fit_transform(data_scaled) # Apply t-SNE for further dimensionality reduction tsne = TSNE(n_components=2, random_state=42) data_tsne = tsne.fit_transform(data_scaled) df = dummy_data features = dummy_data.drop(['Name', 'Condition'], axis=1) columns_of_interest = features.columns.to_list() # Apply K-means on the t-SNE components n_clusters = 10 kmeans = KMeans(n_clusters=n_clusters, random_state=42) labels = kmeans.fit_predict(data_tsne) # Add t-SNE components and cluster labels to the original DataFrame df['TSNE_Component_1'] = data_tsne[:, 0] df['TSNE_Component_2'] = data_tsne[:, 1] df['Cluster'] = labels # Get the centroid coordinates centroids = pd.DataFrame(scaler.inverse_transform(kmeans.cluster_centers_), columns=columns_of_interest) # Display the main features for each centroid for cluster_num in range(n_clusters): centroid_features = centroids.iloc[cluster_num] main_features = centroid_features.abs().sort_values(ascending=False).head(3) # Display top 3 features print(f"Cluster {cluster_num + 1}: Main Features - {main_features.index.tolist()}") # Count the number of users in each cluster cluster_counts = df['Cluster'].value_counts().reset_index() cluster_counts.columns = ['Cluster', 'Number_of_Users'] # Select the top 10 clusters based on the highest number of users top_clusters = cluster_counts.nlargest(10, 'Number_of_Users')['Cluster'].tolist() # Filter the DataFrame for the top clusters df_top_clusters = df[df['Cluster'].isin(top_clusters)]
运行代码时出现错误:
centroids = pd.DataFrame(scaler.inverse_transform(kmeans.cluster_centers_), columns=columns_of_interest) ValueError: operands could not be broadcast together with shapes (10,2) (8,) (10,2)
朋友建议使用其他工具将数据从非线性转为线性降维,但我认为这正是t-SNE的用途,对此存在疑问,希望得到解答。
错误原因分析
scaler基于原始8维特征训练,inverse_transform仅能处理与原始特征维度一致的数据(即8维)。kmeans.cluster_centers_是在t-SNE降维后的2维数据上计算出的簇中心,维度为(10,2),与scaler要求的8维不匹配,因此触发报错。
修复方案
要获取原始特征空间的簇中心,有两种可行方式:
方式一:在原始标准化特征上做K-Means,t-SNE仅用于可视化
将K-Means的拟合对象替换为原始标准化特征data_scaled,t-SNE只负责生成可视化用的低维坐标:
# 修正部分:在原始标准化特征上运行K-Means kmeans = KMeans(n_clusters=n_clusters, random_state=42) labels = kmeans.fit_predict(data_scaled) # t-SNE可视化部分保持不变 df['TSNE_Component_1'] = data_tsne[:, 0] df['TSNE_Component_2'] = data_tsne[:, 1] df['Cluster'] = labels # 现在可正确逆变换得到原始特征空间的簇中心 centroids = pd.DataFrame(scaler.inverse_transform(kmeans.cluster_centers_), columns=columns_of_interest)
方式二:若坚持在t-SNE结果上聚类,基于原始数据计算簇中心
如果一定要在t-SNE降维后的结果上执行K-Means,可通过原始数据中每个簇的样本,计算它们在原始特征空间的均值作为簇中心:
# 修正部分:基于原始特征计算每个簇的中心 centroids = df.groupby('Cluster')[columns_of_interest].mean()
关于t-SNE的疑问解答
t-SNE确实是非线性降维工具,但它的核心定位是数据可视化,而非用于后续聚类的特征空间:
- t-SNE仅保留原始数据的局部相似性,不保留全局结构,其降维空间中簇的距离无法完全代表原始空间的距离。
- t-SNE的变换是不可逆的,无法从低维坐标还原回原始高维特征,这也是之前用scaler逆变换失败的根本原因之一。
如果需要做非线性降维同时保留适合聚类的特征空间,可考虑以下工具:
- UMAP:相比t-SNE,能更好保留全局结构,计算速度更快,且支持从低维近似映射回高维。
- Kernel PCA:通过核函数将数据映射到高维线性空间后做PCA降维,适合处理非线性可分数据,降维结果可直接用于聚类。
内容的提问来源于stack exchange,提问作者Vui Chee Chang
相关产品推荐
相关产品推荐

