You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Wasserstein距离聚类不等长数据时遇AttributeError问题求助

问题解决:Wasserstein距离聚类中的AttributeError错误

错误原因

你遇到的AttributeError: 'list' object has no attribute 'reshape',是因为从data中取出的元素本质是list类型(由于你用np.array包裹了不等长的列表,得到的是dtype=object的数组,每个元素仍是原生list),而你的wasserstein_distance_function函数中直接对其调用了reshape方法,导致报错。

另外,你没有给出wasserstein_distance_function的具体实现,这是核心问题——针对不等长序列的Wasserstein距离需要正确处理输入类型与分布权重。

解决方案

1. 实现正确的Wasserstein距离函数

直接使用scipy.stats提供的现成函数(推荐,避免手动实现的误差),或者在自定义函数中先将输入转为numpy数组:

方法一:使用scipy现成函数

from scipy.stats import wasserstein_distance

def wasserstein_distance_function(x, y):
    # 将输入转为numpy数组,确保后续操作兼容
    x_arr = np.asarray(x)
    y_arr = np.asarray(y)
    # 假设每个样本点的权重相等,直接计算Wasserstein距离
    return wasserstein_distance(x_arr, y_arr)

方法二:自定义最优传输实现(针对离散分布)

如果你需要手动实现基于线性规划的Wasserstein距离:

def wasserstein_distance_function(x, y):
    # 强制转为numpy数组并调整形状
    x_arr = np.asarray(x).reshape(-1, 1)
    y_arr = np.asarray(y).reshape(-1, 1)
    
    # 生成成本矩阵(两点间的绝对距离)
    cost_matrix = np.abs(x_arr - y_arr.T)
    
    # 定义每个样本的权重(均匀分布)
    x_weights = np.ones(len(x_arr)) / len(x_arr)
    y_weights = np.ones(len(y_arr)) / len(y_arr)
    
    # 求解线性分配问题(最优传输路径)
    row_ind, col_ind = linear_sum_assignment(cost_matrix)
    
    # 计算最终的Wasserstein距离
    return cost_matrix[row_ind, col_ind].dot(x_weights)

2. 生成距离矩阵并完成聚类

修正距离函数后,重新生成距离矩阵,再用层次聚类完成分类:

import numpy as np
from scipy.optimize import linear_sum_assignment
from scipy.cluster.hierarchy import linkage, fcluster
from sklearn.cluster import AgglomerativeClustering
from scipy.stats import wasserstein_distance

# 定义数据
data = np.array([[5, 2, 2],[3, 6],[1, 6, 2],[7, 2],[7, 2], [6, 1, 3],[8, 1], [2, 4, 7, 3]])

# 实现Wasserstein距离函数
def wasserstein_distance_function(x, y):
    x_arr = np.asarray(x)
    y_arr = np.asarray(y)
    return wasserstein_distance(x_arr, y_arr)

# 生成对称距离矩阵
n = len(data)
distance_matrix = np.zeros((n, n))
for i in range(n):
    for j in range(i, n):
        dist = wasserstein_distance_function(data[i], data[j])
        distance_matrix[i][j] = dist
        distance_matrix[j][i] = dist

# 方法一:使用sklearn进行层次聚类
cluster = AgglomerativeClustering(n_clusters=3, metric='precomputed', linkage='average')
labels = cluster.fit_predict(distance_matrix)
print("聚类结果:", labels)

# 方法二:使用scipy层次聚类
# 将距离矩阵转为condensed形式
from scipy.spatial.distance import squareform
condensed_dist = squareform(distance_matrix)
linkage_matrix = linkage(condensed_dist, method='ward')
labels_scipy = fcluster(linkage_matrix, t=3, criterion='maxclust')
print("Scipy聚类结果:", labels_scipy)

关键注意点

  • 当处理不等长序列时,np.array会生成dtype=object的数组,每个元素仍是list,必须显式转为numpy数组后再进行数值操作。
  • 聚类时,sklearn.AgglomerativeClustering支持直接传入预计算的距离矩阵(需设置metric='precomputed'),而scipy.linkage需要将距离矩阵转为condensed的一维形式。

内容的提问来源于stack exchange,提问作者aam

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.12 23:10:25