You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用向量化实现NumPy数组的分组累积计数?

用NumPy实现按标签的累积计数需求

原始输入数组

import numpy as np
y = np.array([
    [ 0],
    [ 0],
    [ 2],
    [ 1],
    [ 0],
    [ 1],
    [ 3],
    [-1],
])

期望输出

为原数组添加新列,记录每个标签值的累积出现次数:

np.array([
    [ 0, 1],
    [ 0, 2],
    [ 2, 1],
    [ 1, 1],
    [ 0, 3],
    [ 1, 2],
    [ 3, 1],
    [-1, 1],
])

已实现的Pandas方案

import pandas as pd
y_pd = pd.DataFrame(y, columns=['LABEL'])
y_pd = pd.concat([
    y_pd, 
    y_pd.groupby('LABEL').cumcount().to_frame().rename(columns = {0:'cumcounts'}) +1
], axis=1)

当前的NumPy循环实现

y_np = np.hstack([y, y])
for label in np.unique(y_np):
    slice_length = (y_np[:, -2]==label).sum()
    y_np[y_np[:, -2]==label, -1] = range(1, slice_length+1)

向量化NumPy解决方案

以下是完全向量化的实现,无循环,适合大规模数组且保留原始顺序:

# 扁平化原始数组以便处理
y_flat = y.flatten()

# 获取标签排序后的索引,将相同标签归组
sorted_idx = np.argsort(y_flat)
sorted_labels = y_flat[sorted_idx]

# 初始化计数数组,所有位置先设为1
counts = np.ones_like(y_flat)

# 找到标签发生变化的位置(跳过第一个元素)
diff_mask = sorted_labels[1:] != sorted_labels[:-1]

# 对非标签变化的位置,计数在前一个同标签位置的基础上加1;变化位置保持为1
counts[sorted_idx[1:]] = np.where(diff_mask, 1, counts[sorted_idx[:-1]] + 1)

# 将计数数组重塑为列,与原数组合并得到最终结果
result = np.hstack([y, counts.reshape(-1, 1)])

方案说明

  1. 通过argsort将相同标签的元素索引归为一组,便于批量计算累积计数
  2. 利用diff标记不同标签的边界,在边界位置重置计数为1,非边界位置累加计数
  3. 通过排序索引将计数还原回原始数组的顺序,确保记录顺序不变

内容的提问来源于stack exchange,提问作者lemon

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.02 11:32:48