如何创建按Over分组统计length/type唯一实例计数的列
问题
需要创建名为length/type_count的列,统计每个Over内length/type变量的唯一实例数量。每6个deliveries_faced开启一个新Over,计数需在每个Over开始时(每6次投递后)重置。示例数据如下:
| Batfast_id | Session_no | Event_name | Overs | deliveries_faced | length/type | length/type_count |
|---|---|---|---|---|---|---|
| player1 | 1 | ful | 1 | 0 | S_S_Y | 1 |
| player1 | 1 | ful | 1 | 1 | S_S_Y | 1 |
| player1 | 1 | ful | 1 | 2 | M_S_Y | 2 |
| player1 | 1 | ful | 1 | 3 | M_OS_F | 3 |
| player1 | 1 | ful | 1 | 4 | M_S_Y | 3 |
| player1 | 1 | ful | 1 | 5 | F_LS_Y | 4 |
| player1 | 1 | ful | 2 | 6 | M_S_F | 1 |
| player1 | 1 | ful | 2 | 7 | S_S_ES | 2 |
| player1 | 1 | ful | 2 | 8 | S_S_ES | 2 |
| player1 | 1 | ful | 2 | 9 | S_S_ES | 2 |
| player1 | 1 | ful | 2 | 10 | F_S_Y | 3 |
| player1 | 1 | ful | 2 | 11 | F_S_Y | 3 |
注:Batfast_id、Session_no、Event_name、Overs、deliveries_faced为数据集的索引,同时也作为列存在。
解决方案
可以用Pandas实现需求,核心是按Over分组后,对length/type做累计唯一值计数:
import pandas as pd # 重置索引(如果原数据集把指定列设为索引) df = df.reset_index() # 定义函数计算累计唯一值数量 def cumulative_unique_count(series): seen = set() count_list = [] for val in series: seen.add(val) count_list.append(len(seen)) return count_list # 按分组计算并生成目标列 df['length/type_count'] = df.groupby( ['Batfast_id', 'Session_no', 'Event_name', 'Overs'] )['length/type'].transform(cumulative_unique_count) # 恢复原索引(可选) df = df.set_index(['Batfast_id', 'Session_no', 'Event_name', 'Overs', 'deliveries_faced'])
代码说明
- 先重置索引,确保用于分组的列能被正常访问;
- 自定义函数
cumulative_unique_count遍历每组内的length/type值,用集合记录已出现的唯一值,每次迭代时记录集合的长度,即当前累计的唯一实例数; - 通过
groupby+transform将计算结果映射回原数据集,保证每个行对应正确的计数; - 最后可按需恢复原索引结构。
内容的提问来源于stack exchange,提问作者Yoseph Ismail
相关产品推荐
相关产品推荐

