如何在Pandas中基于另一列生成含元素统计字典的新列?
解决方案:生成campaign_type元素计数字典列
这个需求很容易实现,我们可以用Python的pandas库配合collections.Counter来快速生成统计字典列,下面是完整的步骤和代码示例:
1. 准备依赖和示例数据
首先导入需要的库,然后构造你提供的示例DataFrame:
import pandas as pd from collections import Counter # 构造示例数据 data = { 'date': ['2019-01-01', '2019-01-01', '2019-02-01', '2019-02-01'], 'product': ['Dell', 'Apple', 'Dell', 'Apple'], 'campaign_type': [['call', 'email', 'call'], ['fax', 'fax', 'visit', 'visit'], ['call', 'fax', 'call'], ['email', 'email', 'visit']], 'total_monthly_sale': [5, 4, 6, 7] } df = pd.DataFrame(data)
2. 生成campaign_dict列
用apply方法遍历campaign_type列的每个列表,通过Counter统计元素出现次数,再转成普通字典:
# 新增统计字典列 df['campaign_dict'] = df['campaign_type'].apply(lambda x: dict(Counter(x)))
3. 查看结果
执行完上面的代码后,打印DataFrame就能得到你想要的输出:
print(df)
预期输出表格:
| date | product | campaign_type | total_monthly_sale | campaign_dict |
|---|---|---|---|---|
| 2019-01-01 | Dell | ['call', 'email', 'call'] | 5 | {'call': 2, 'email': 1} |
| 2019-01-01 | Apple | ['fax', 'fax', 'visit', 'visit'] | 4 | {'fax': 2, 'visit': 2} |
| 2019-02-01 | Dell | ['call', 'fax', 'call'] | 6 | {'call': 2, 'fax': 1} |
| 2019-02-01 | Apple | ['email', 'email', 'visit'] | 7 | {'email': 2, 'visit': 1} |
额外提示:处理字符串形式的列表
如果你的campaign_type列存储的是字符串格式的列表(比如像"[call, email, call]"这样的字符串,而不是真正的Python列表),需要先把它转换成列表再统计:
import ast # 将字符串形式的列表转为真正的列表 df['campaign_type'] = df['campaign_type'].apply(ast.literal_eval) # 再执行统计操作 df['campaign_dict'] = df['campaign_type'].apply(lambda x: dict(Counter(x)))
内容的提问来源于stack exchange,提问作者StackUser
相关产品推荐
相关产品推荐

