如何用Python/Pandas统计不区分顺序的宝可梦对战类型组合计数?
解决宝可梦对战组合统计的顺序问题
要解决bug vs dark和dark vs bug被视为不同组合的问题,核心思路是统一每对战组合的类型顺序,让不管谁先谁后,相同的类型对都被归为同一组。下面是几种实用的实现方式:
方法1:生成排序后的组合列(直观易读)
先给DataFrame新增一列,存储每行两个宝可梦类型排序后的元组,再按这个新列分组统计:
# 对每行的两个类型进行排序,生成统一的组合 combat['type_combination'] = combat[['First_pokemon_type', 'Second_pokemon_type']].apply( lambda row: tuple(sorted(row)), axis=1 ) # 按新的组合列统计数量 result = combat.groupby('type_combination', as_index=False).agg( battle_count=('Winner_pokemon_type', 'count') )
方法2:直接在groupby中处理(无需新增列)
如果不想额外增加列,可以直接在groupby的by参数里完成排序逻辑:
result = combat.groupby( by=combat[['First_pokemon_type', 'Second_pokemon_type']].apply( lambda row: tuple(sorted(row)), axis=1 ), as_index=False ).agg(battle_count=('Winner_pokemon_type', 'count'))
方法3:用numpy提升大数据量下的效率
如果你的DataFrame数据量很大,用numpy的sort函数会比apply更快:
import numpy as np # 对类型列按行排序 sorted_types = np.sort(combat[['First_pokemon_type', 'Second_pokemon_type']], axis=1) # 将排序后的结果拆分为两列 combat['sorted_type1'] = sorted_types[:, 0] combat['sorted_type2'] = sorted_types[:, 1] # 按排序后的两列分组统计 result = combat.groupby(['sorted_type1', 'sorted_type2'], as_index=False).agg( battle_count=('Winner_pokemon_type', 'count') )
这几种方法都能实现把bug vs dark和dark vs bug合并为同一组合统计,你可以根据自己的代码习惯和数据规模选择合适的方式。
内容的提问来源于stack exchange,提问作者Andres Uribe
相关产品推荐
相关产品推荐

