如何为数组创建经验分布?验证wasserstein_distance输入有效性
关于Scipy Wasserstein距离的经验分布输入问题
Scipy的wasserstein_distance函数要求输入为**(经验)分布中的观测值**,现有以下两组范围在-4到8之间的数组:
import numpy as np x = np.array([0.12,-1.29,-3.23,-3.21,-0.13, 1.52, 4.45, 6.45, 5.17, 0.11, 3.48, 5.98, 7.55]) y = np.array([3.54, 2.42,-4.43,-3.76, 0.43, 0.45, 2.56, 7.61, 4.47, 1.36, 2.34, 7.78, 7.13])
问题说明
你尝试用statsmodels的ECDF构造经验分布,代码如下:
from statsmodels.distributions.empirical_distribution import ECDF ecdf_x = ECDF(x) x_ecdf = ecdf_y.y # 这里存在笔误,应为ecdf_x.y ecdf_y = ECDF(y) y_ecdf = ecdf_y.y wasserstein_distance(x_ecdf, y_ecdf)
想知道x_ecdf和y_ecdf是否是wasserstein_distance的有效输入?
答案:无效
原因:
wasserstein_distance的设计逻辑是直接接收原始样本观测值(也就是你手里的x和y数组),它会自动基于这些样本计算经验分布,进而计算距离,不需要手动传入ECDF的输出。- statsmodels的
ECDF.y返回的是累积概率值序列(取值在0到1之间),这完全不是函数要求的“观测值”,传入后计算出的结果没有任何统计意义。
正确用法:
直接将原始数组传入函数即可:
from scipy.stats import wasserstein_distance import numpy as np x = np.array([0.12,-1.29,-3.23,-3.21,-0.13, 1.52, 4.45, 6.45, 5.17, 0.11, 3.48, 5.98, 7.55]) y = np.array([3.54, 2.42,-4.43,-3.76, 0.43, 0.45, 2.56, 7.61, 4.47, 1.36, 2.34, 7.78, 7.13]) # 直接传入原始观测值 distance = wasserstein_distance(x, y) print(distance)
进阶用法(手动指定分位点与权重):
如果需要手动构造经验分布的分位点和对应权重,函数也支持这种输入方式(结果和直接传原始数组一致):
x_sorted = np.sort(x) y_sorted = np.sort(y) # 每个样本的权重为1/样本总数(均匀分布) x_weights = np.full(len(x), 1/len(x)) y_weights = np.full(len(y), 1/len(y)) distance = wasserstein_distance(x_sorted, y_sorted, x_weights, y_weights)
内容的提问来源于stack exchange,提问作者HappyPy
相关产品推荐
相关产品推荐

