You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何为数组创建经验分布?验证wasserstein_distance输入有效性

关于Scipy Wasserstein距离的经验分布输入问题

Scipy的wasserstein_distance函数要求输入为**(经验)分布中的观测值**,现有以下两组范围在-4到8之间的数组:

import numpy as np
x = np.array([0.12,-1.29,-3.23,-3.21,-0.13, 1.52, 4.45, 6.45, 5.17, 0.11, 3.48, 5.98, 7.55])
y = np.array([3.54, 2.42,-4.43,-3.76, 0.43, 0.45, 2.56, 7.61, 4.47, 1.36, 2.34, 7.78, 7.13])

问题说明

你尝试用statsmodels的ECDF构造经验分布,代码如下:

from statsmodels.distributions.empirical_distribution import ECDF

ecdf_x = ECDF(x)
x_ecdf = ecdf_y.y  # 这里存在笔误,应为ecdf_x.y

ecdf_y = ECDF(y)
y_ecdf = ecdf_y.y

wasserstein_distance(x_ecdf, y_ecdf)

想知道x_ecdf和y_ecdf是否是wasserstein_distance的有效输入?


答案:无效

原因:

  1. wasserstein_distance的设计逻辑是直接接收原始样本观测值(也就是你手里的x和y数组),它会自动基于这些样本计算经验分布,进而计算距离,不需要手动传入ECDF的输出。
  2. statsmodels的ECDF.y返回的是累积概率值序列(取值在0到1之间),这完全不是函数要求的“观测值”,传入后计算出的结果没有任何统计意义。

正确用法:

直接将原始数组传入函数即可:

from scipy.stats import wasserstein_distance
import numpy as np

x = np.array([0.12,-1.29,-3.23,-3.21,-0.13, 1.52, 4.45, 6.45, 5.17, 0.11, 3.48, 5.98, 7.55])
y = np.array([3.54, 2.42,-4.43,-3.76, 0.43, 0.45, 2.56, 7.61, 4.47, 1.36, 2.34, 7.78, 7.13])

# 直接传入原始观测值
distance = wasserstein_distance(x, y)
print(distance)

进阶用法(手动指定分位点与权重):

如果需要手动构造经验分布的分位点和对应权重,函数也支持这种输入方式(结果和直接传原始数组一致):

x_sorted = np.sort(x)
y_sorted = np.sort(y)
# 每个样本的权重为1/样本总数(均匀分布)
x_weights = np.full(len(x), 1/len(x))
y_weights = np.full(len(y), 1/len(y))

distance = wasserstein_distance(x_sorted, y_sorted, x_weights, y_weights)

内容的提问来源于stack exchange,提问作者HappyPy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.03 12:02:42