如何更高效地基于x、y区间关联的z值平均值填充二维Numpy数组?
更高效的Numpy实现方案
当然有更高效的numpy原生方法来替代嵌套循环!你的思路完全没问题,但numpy内置的统计函数能帮你大幅简化代码,还能提升运行效率——尤其是当你的数据量变大的时候。
方法一:使用np.histogram2d(推荐)
np.histogram2d可以一次性统计每个x-y单元格内的点数,以及这些点对应的z值总和,我们只需要用总和除以点数就能得到平均值,全程不需要手动循环。
import numpy as np import matplotlib.pyplot as plt x = [37.59390426045407, 38.00530354847739, 38.28412244348653, 38.74871247986305, 38.73175910429809, 38.869008864244016, 39.188234404976555, 39.92835838352555, 40.881394113153334, 41.686136269465884] y = [0.1305391767832006, 0.13764519613447768, 0.14573326951792354, 0.15090729309032114, 0.16355823707239897, 0.17327106424274763, 0.17749746339532224, 0.17310384614773594, 0.16545780437882962, 0.1604752704890856] z = [0.05738534353865021, 0.012572155256903583, -0.021709582561809437, -0.11191337750722108, -0.07931921785775153, -0.06241610118871843, 0.014216349927058225, 0.042002641153291886, -0.029354425271534645, 0.061894011359833856] n = 5 # 转换为numpy数组方便后续操作 x_np = np.array(x) y_np = np.array(y) z_np = np.array(z) # 生成x和y方向的分箱边界,确保分成n个等宽单元格 x_bins = np.linspace(x_np.min(), x_np.max(), n + 1) y_bins = np.linspace(y_np.min(), y_np.max(), n + 1) # 统计每个单元格内的点数,以及z值的加权总和 cell_counts, _, _ = np.histogram2d(x_np, y_np, bins=(x_bins, y_bins)) cell_z_sum, _, _ = np.histogram2d(x_np, y_np, bins=(x_bins, y_bins), weights=z_np) # 计算平均值,用where参数避免除以0的情况 image = np.divide(cell_z_sum, cell_counts, where=cell_counts != 0) # 没有数据的单元格保持为0(你也可以改成np.nan或者其他值) image[cell_counts == 0] = 0 # 绘图时注意转置数组并设置origin='lower',让坐标系和x-y实际方向一致 fig, ax = plt.subplots() ax.imshow(image.T, cmap=plt.cm.gray, interpolation='nearest', origin='lower') plt.show()
方法二:使用np.digitize + np.bincount
如果你需要更灵活的索引控制,可以先通过np.digitize找到每个数据点对应的单元格索引,再用np.bincount统计每个单元格的z总和和点数:
# 接上面的x_np, y_np, z_np和n、bins定义 x_indices = np.digitize(x_np, x_bins) - 1 # 把索引调整为从0开始 y_indices = np.digitize(y_np, y_bins) - 1 # 将二维索引转换为线性索引,方便bincount统计 linear_indices = x_indices * n + y_indices # 统计每个线性索引对应的z总和和点数 sum_z = np.bincount(linear_indices, weights=z_np, minlength=n*n) counts = np.bincount(linear_indices, minlength=n*n) # 转换回二维数组并计算平均值 image = sum_z.reshape(n, n) / counts.reshape(n, n) image[counts == 0] = 0
为什么这两种方法更好?
- 效率更高:numpy的内置函数是用C实现的,比Python嵌套循环快得多,数据量越大优势越明显
- 代码更简洁:不需要手动写两层循环,逻辑更清晰,可读性更强
- 更安全:
np.divide的where参数可以避免除以0的错误,处理空单元格更稳妥
内容的提问来源于stack exchange,提问作者Alex
相关产品推荐
相关产品推荐

