如何在Pandas中基于分组与其他行列值生成term统计新列?
Pandas实现按月份统计包含当前term值的其他行数需求
问题描述
给定以下DataFrame:
month term 0 Jun-22 one 1 Jun-22 one two 2 Jul-22 one 3 Jul-22 three 4 Jul-22 three four 5 Jul-22 three four five
需要添加一列term_count,用于统计对应月份中包含当前行term值的其他行数(排除当前行本身)。
期望输出
month term term_count 0 Jun-22 one 1 1 Jun-22 one two 0 2 Jul-22 one 0 3 Jul-22 three 2 4 Jul-22 three four 1 5 Jul-22 three four five 0
示例说明
- 第0行的
term为'one',Jun-22月份的其他1行(第1行)包含该值,故term_count为1; - 第1行的
term为'one two',Jun-22月份的其他行无包含该值的,故term_count为0; - 第3行的
term为'three',Jul-22月份的其他2行(第4、5行)包含该值,故term_count为2。
解决方案
可以通过按月份分组+组内遍历统计的方式实现,具体代码如下:
步骤1:创建示例DataFrame
import pandas as pd data = { 'month': ['Jun-22', 'Jun-22', 'Jul-22', 'Jul-22', 'Jul-22', 'Jul-22'], 'term': ['one', 'one two', 'one', 'three', 'three four', 'three four five'] } df = pd.DataFrame(data)
步骤2:实现统计逻辑
def calculate_term_count(group): # 获取组内所有term的列表 term_list = group['term'].tolist() # 对每个term,统计组内其他行中包含该term的数量 group['term_count'] = [ sum(1 for t in term_list if term in t) - 1 # 减1排除当前行自身 for term in term_list ] return group # 按month分组应用统计函数 df = df.groupby('month', group_keys=False).apply(calculate_term_count)
逻辑说明
- 按
month列分组,确保只在同一月份内统计; - 对每个分组,遍历每一行的
term:- 先统计组内所有行(包括当前行)中包含该
term的数量; - 减去1,排除当前行自身的计数;
- 先统计组内所有行(包括当前行)中包含该
- 将统计结果赋值给新列
term_count。
执行后即可得到符合要求的DataFrame。
内容的提问来源于stack exchange,提问作者Adam
相关产品推荐
相关产品推荐

