You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用scipy的binned_statistic_2d实现等量样本的2D直方图

Hey there! 我懂你想在二维数据集上实现等样本量分箱(每个区间里的样本数尽量相等),然后用scipy.stats.binned_statistic_2d来计算统计量的需求~从一维代码扩展到二维确实容易卡壳,我来一步步帮你拆解和实现。

先理清楚一维等样本分箱的逻辑

你给出的一维代码核心思路非常巧妙:把数据排序后,按照累积样本数量的分位点来确定分箱边界,这样每个区间里的样本数自然相等(或近似相等)。np.interp在这里的作用是,把“样本数的分位点”映射回“实际数据的数值边界”,完美解决了等样本分箱的问题。

扩展到二维:轴对齐的等样本分箱

最直接且和你现有代码衔接最紧密的思路是:分别对x和y维度做一维等样本分箱,然后把两个维度的边界组合成二维网格。这种方法能直接适配binned_statistic_2d的参数要求。

完整实现代码

import numpy as np
import scipy.stats as stats
import matplotlib.pyplot as plt

# 复用你的一维等样本分箱函数
def histedges_equalN(data, nbin):
    npt = len(data)
    return np.interp(np.linspace(0, npt, nbin + 1), np.arange(npt), np.sort(data))

# 生成测试二维数据(带相关性,更贴近真实场景)
np.random.seed(42)
x = np.random.randn(1000)
y = np.random.randn(1000) + 0.5 * x  # 让x和y有线性相关性

# 设定分箱数:x和y各分10箱,总共100个二维区间
n_bins_x = 10
n_bins_y = 10

# 生成x和y的等样本分箱边界
bins_x = histedges_equalN(x, n_bins_x)
bins_y = histedges_equalN(y, n_bins_y)

# 调用binned_statistic_2d计算统计量(这里以计数为例,你可以换成'mean'/'sum'等)
stat, x_edges, y_edges, binnumber = stats.binned_statistic_2d(
    x, y, values=None, statistic='count',
    bins=[bins_x, bins_y], expand_binnumbers=True
)

# 查看每个二维箱的样本数(应该接近10,因为1000/100=10)
print("每个二维箱的样本数:")
print(stat.astype(int))

# 可视化分箱效果
plt.figure(figsize=(8, 6))
plt.scatter(x, y, c=binnumber[0] * n_bins_y + binnumber[1], cmap='tab20', alpha=0.6)
plt.hlines(bins_y, xmin=x.min(), xmax=x.max(), color='gray', linestyle='--', alpha=0.5)
plt.vlines(bins_x, ymin=y.min(), ymax=y.max(), color='gray', linestyle='--', alpha=0.5)
plt.xlabel('X')
plt.ylabel('Y')
plt.title('二维轴对齐等样本分箱')
plt.colorbar(label='箱编号')
plt.show()

关键细节解释

  1. 分箱边界生成:histedges_equalN对x和y分别处理,确保每个x区间和y区间的样本数相等,组合后二维箱的样本数会尽可能接近平均值。
  2. binned_statistic_2d参数:bins=[bins_x, bins_y]指定两个维度的边界,expand_binnumbers=True会返回每个样本对应的x/y箱编号,方便后续可视化或进一步处理。
  3. 样本数差异:如果总样本数不能被分箱总数整除,少数箱的样本数会多1或少1,这是正常现象,不会影响整体的等样本效果。
进阶:严格等样本的非轴对齐分箱

如果需要每个二维箱的样本数完全相等,轴对齐的方法可能做不到(因为数据的联合分布可能不均匀)。这时候可以用K-Means聚类来实现:

from sklearn.cluster import KMeans

n_total_bins = 100  # 总共100个箱,每个箱10个样本
data_2d = np.column_stack((x, y))

# K-Means聚类,确保每个簇样本数相等
kmeans = KMeans(n_clusters=n_total_bins, random_state=42, n_init='auto')
labels = kmeans.fit_predict(data_2d)

# 验证每个簇的样本数
unique, counts = np.unique(labels, return_counts=True)
print("每个簇的样本数:", counts)  # 输出全是10,严格相等

# 可视化聚类结果
plt.figure(figsize=(8, 6))
plt.scatter(x, y, c=labels, cmap='tab20', alpha=0.6)
plt.scatter(kmeans.cluster_centers_[:, 0], kmeans.cluster_centers_[:, 1], 
            c='red', marker='x', s=100, label='簇中心')
plt.xlabel('X')
plt.ylabel('Y')
plt.title('二维严格等样本分箱(非轴对齐)')
plt.legend()
plt.show()

注意点

  • 这种方法的箱是不规则形状的,不适合需要轴对齐边界的场景,但能保证每个箱的样本数完全一致。
  • 需要安装scikit-learn库(执行pip install scikit-learn即可)。

内容的提问来源于stack exchange,提问作者elizevin

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 11:52:42