如何用Pandas统计文本事件类型并转换为国家-年度数据
解决方案:统计指定目标类型的年度事件数并合并伤亡汇总
方法一:创建标记列后一次性聚合(直观易懂)
适合目标类型数量较少的场景,先为每个指定目标类型生成标记列,再统一分组计算:
- 生成目标类型标记列(注意名称需与数据中
targettype_txt的取值完全匹配,包括大小写、空格):
# 替换为你实际需要统计的5类目标 target_list = ['government building', 'military', 'police', 'civilian', 'business'] # 为每个目标类型添加标记列:匹配则为1,否则为0 for target in target_list: # 把列名中的空格替换为下划线,避免后续使用报错 col_name = target.replace(' ', '_') cdf[col_name] = (cdf['targettype_txt'] == target).astype(int)
- 执行分组聚合,同时完成伤亡汇总和目标事件计数:
final_df = cdf.groupby(['country', 'country_txt', 'iyear']).agg( nkill=('nkill', 'sum'), nwound=('nwound', 'sum'), nhostkid=('nhostkid', 'sum'), government_building=('government_building', 'sum'), military=('military', 'sum'), police=('police', 'sum'), civilian=('civilian', 'sum'), business=('business', 'sum') )
方法二:用聚合字典批量处理(高效简洁)
如果目标类型较多,用字典定义聚合规则可以减少重复代码:
target_list = ['government building', 'military', 'police', 'civilian', 'business'] # 基础聚合规则:伤亡、人质的求和 agg_rules = { 'nkill': 'sum', 'nwound': 'sum', 'nhostkid': 'sum' } # 为每个目标类型添加计数规则 for target in target_list: col_name = target.replace(' ', '_') agg_rules[col_name] = ('targettype_txt', lambda x: (x == target).sum()) # 执行分组聚合 final_df = cdf.groupby(['country', 'country_txt', 'iyear']).agg(**agg_rules)
方法三:合并已有结果与目标计数
如果想保留你已生成的df2,可以单独统计目标事件数后再合并:
target_list = ['government building', 'military', 'police', 'civilian', 'business'] # 统计每个国家-年度的各目标类型事件数,缺失值填充为0 target_counts = cdf.groupby(['country', 'country_txt', 'iyear'])['targettype_txt'].value_counts().unstack(fill_value=0)[target_list] # 重命名列(可选) target_counts.columns = [col.replace(' ', '_') for col in target_counts.columns] # 合并已有汇总结果和目标计数 final_df = df2.join(target_counts, how='outer')
关键注意事项
- 务必确认目标类型名称与
targettype_txt字段中的取值完全一致,可以通过print(cdf['targettype_txt'].unique())查看所有可能的取值,避免因名称不匹配导致统计为0。 - 若数据存在缺失值,可根据需求在聚合或合并时添加
dropna参数处理。
内容的提问来源于stack exchange,提问作者taraamcl
相关产品推荐
相关产品推荐

