如何在Pandas中为DataFrame的cluster列各唯一值生成累计求和列?
实现Pandas按类别生成累计计数列
问题背景
现有如下Pandas DataFrame:
import pandas as pd import random random.seed(42) df = pd.DataFrame({'index': list(range(0,10)), 'cluster': [random.choice(['S', 'C']) for l in range(0,10)]})
生成的DataFrame结构为:
index cluster 0 0 S 1 1 S 2 2 C 3 3 S 4 4 S 5 5 S 6 6 S 7 7 S 8 8 C 9 9 S
需求
为cluster列的每个唯一值创建新列,存储该值出现次数的累计求和结果,最终得到如下结构的DataFrame:
index cluster cumulative_S cumulative_C 0 0 S 1 0 1 1 S 2 0 2 2 C 2 1 3 3 S 3 1 4 4 S 4 1 5 5 S 5 1 6 6 S 6 1 7 7 S 7 1 8 8 C 7 2 9 9 S 8 2
解决方案
可以通过以下步骤实现:
- 生成哑变量矩阵,将
cluster列的类别转为0-1标记的独立列 - 对哑变量矩阵做累计求和,得到每个类别的累计出现次数
- 重命名累计列并合并回原DataFrame
完整代码如下:
import pandas as pd import random random.seed(42) # 生成原始DataFrame df = pd.DataFrame({'index': list(range(0,10)), 'cluster': [random.choice(['S', 'C']) for l in range(0,10)]}) # 生成哑变量并计算累计求和 cumulative_counts = pd.get_dummies(df['cluster']).cumsum() # 重命名列,添加指定前缀 cumulative_counts = cumulative_counts.rename(columns=lambda x: f'cumulative_{x}') # 合并到原DataFrame result_df = pd.concat([df, cumulative_counts], axis=1) print(result_df)
代码说明
pd.get_dummies(df['cluster']):把cluster的每个类别转为单独列,匹配类别的位置标记为1,其余为0.cumsum():对每列做逐行累计求和,得到从第一行到当前行的类别累计出现次数rename(columns=lambda x: f'cumulative_{x}'):给累计计数列添加需求指定的前缀,统一命名格式pd.concat([df, cumulative_counts], axis=1):将原数据和累计计数列横向合并,得到最终结果
内容的提问来源于stack exchange,提问作者quant
相关产品推荐
相关产品推荐

