You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于索引数组以向量化方式对值数组求和?

向量化实现按索引分组求和

我有一个值数组:

values = np.array([0.0, 1.0, 2.0, 3.0, 4.0])

以及一个索引数组:

indices = np.array([0,1,0,2,2])

想要用向量化方式实现以下循环代码的功能:对indices中每个唯一索引对应的值求和:

sums = np.zeros(np.max(indices)+1)
for index, value in zip(indices, values):
    sums[index] += value

额外要求:最好支持values(以及对应的sums)为多维数组。


解决方案与基准测试

我对几种可行方案做了基准测试,测试代码如下:

import numpy as np
import time
import pandas as pd


values = np.arange(1_000_000, dtype=float)
rng = np.random.default_rng(0)
indices = rng.integers(0, 1000, size=1_000_000)


N = 100


now = time.time_ns()
for _ in range(N):
    sums = np.bincount(indices, weights=values, minlength=1000)
print(f"np.bincount: {(time.time_ns() - now) * 1e-6 / N:.3f} ms")


now = time.time_ns()
for _ in range(N):
    sums = np.zeros(1 + np.amax(indices), dtype=values.dtype)
    np.add.at(sums, indices, values)
print(f"np.add.at: {(time.time_ns() - now) * 1e-6 / N:.3f} ms")


now = time.time_ns()
for _ in range(N):
    pd.Series(values).groupby(indices).sum().values
print(f"pd.groupby: {(time.time_ns() - now) * 1e-6 / N:.3f} ms")


now = time.time_ns()
for _ in range(N):
    sums = np.zeros(np.max(indices)+1)
    for index, value in zip(indices, values):
        sums[index] += value
print(f"Loop: {(time.time_ns() - now) * 1e-6 / N:.3f} ms")

测试结果:

np.bincount: 1.129 ms
np.add.at: 0.763 ms
pd.groupby: 5.215 ms
Loop: 196.633 ms

方案说明

  • np.bincount:适合非负整数索引场景,通过weights参数传入values即可直接求和,minlength可指定结果数组长度,但仅支持一维values。
  • np.add.at:支持多维values,通过at方法实现指定索引位置的原地累加,相比循环效率提升极大,灵活性更强。
  • pandas.groupby:借助pandas分组求和功能实现,代码简洁但性能弱于numpy原生方法。

内容的提问来源于stack exchange,提问作者Padix Key

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.16 18:05:06