You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Numpy按标签对二维数组元素求和

按标签对Numpy二维数据点求和的Python风格实现

给定二维数据矩阵X:

X = np.array(
    [
        [6, 1], # row_0
        [4, 4], # row_1
        [8, 4], # row_2
        [6, 3], # row_..
        [5, 8],
        [7, 9]  # row_5
    ]
)

以及对应的标签数组labels:

labels = np.array([1, 0, 2, 1, 2, 0])

目前通过循环实现按标签求和:

cum_sum = np.zeros((3, 2))
for i, label in enumerate(labels):
    cum_sum[label] += X[i]

得到结果:

[[11. 13.]
 [12.  4.]
 [13. 12.]]

以下是几种更符合Python风格、高效的Numpy解决方案:

方法1:使用np.bincount(推荐)

np.bincount可按标签统计频次,扩展到二维时,对每一列分别应用该函数后转置即可得到结果:

import numpy as np

X = np.array([[6,1],[4,4],[8,4],[6,3],[5,8],[7,9]])
labels = np.array([1,0,2,1,2,0])

cum_sum = np.array([np.bincount(labels, weights=X[:, col]) for col in range(X.shape[1])]).T
print(cum_sum)

输出:

[[11. 13.]
 [12.  4.]
 [13. 12.]]

该方法完全向量化,无手动循环,数据量越大,效率优势越明显。

方法2:使用np.add.at

np.add.at是Numpy专门针对按索引批量累加场景设计的函数,能避免普通广播的重复索引问题:

import numpy as np

X = np.array([[6,1],[4,4],[8,4],[6,3],[5,8],[7,9]])
labels = np.array([1,0,2,1,2,0])

cum_sum = np.zeros((3, 2))
np.add.at(cum_sum, labels, X)
print(cum_sum)

输出与预期一致,实现简洁且效率优于手动循环。

方法3:分组排序求和(参考)

先按标签排序,再拆分分组求和,步骤稍多,适合理解分组逻辑:

import numpy as np

X = np.array([[6,1],[4,4],[8,4],[6,3],[5,8],[7,9]])
labels = np.array([1,0,2,1,2,0])

# 按标签排序
sorted_indices = np.argsort(labels)
sorted_X = X[sorted_indices]
sorted_labels = labels[sorted_indices]

# 定位标签分界点
split_points = np.where(np.diff(sorted_labels))[0] + 1
# 拆分分组并求和
cum_sum = np.array([np.sum(group, axis=0) for group in np.split(sorted_X, split_points)])
print(cum_sum)

内容的提问来源于stack exchange,提问作者Chiel

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.16 09:50:17