如何将宝可梦数据的type1、type2统计结果合并为单个DataFrame?
宝可梦类型统计的正确实现方法
一、修复你现有的合并问题
你用pd.concat得到重复单列,是因为默认按行(axis=0)拼接了两个统计结果。要得到两列结构,得先把两个统计结果的列名区分开,再按类型索引对齐合并:
假设你之前的统计代码是生成带列名的DataFrame:
df_type1 = df['type1'].value_counts().reset_index(name='type1_count') df_type2 = df['type2'].value_counts().reset_index(name='type2_count')
改成按类型列合并:
# 按类型列全量合并,空值填充为0 result = pd.merge(df_type1, df_type2, on='index', how='outer').fillna(0) # 重命名类型列 result = result.rename(columns={'index': 'type'}) # 把计数转成整数格式 result[['type1_count', 'type2_count']] = result[['type1_count', 'type2_count']].astype(int)
如果你的统计结果是用索引存类型的Series:
s_type1 = df['type1'].value_counts() s_type2 = df['type2'].value_counts() # 按列拼接并自动对齐索引,空值填0 result = pd.concat([s_type1, s_type2], axis=1, keys=['type1_count', 'type2_count']).fillna(0).astype(int) # 可选:把索引转成单独的类型列 result = result.reset_index().rename(columns={'index': 'type'})
二、更高效的一站式统计方案
不用分开统计再合并,直接一步生成目标DataFrame:
# 直接对两列做value_counts,按列拼接并填充空值 type_stats = pd.DataFrame({ 'type1_count': df['type1'].value_counts(), 'type2_count': df['type2'].value_counts() }).fillna(0).astype(int).reset_index().rename(columns={'index': 'type'})
如果需要扩展到更多类型列,用melt转长格式后分组统计更灵活:
# 把type1、type2转成长格式,过滤空值 melted = df.melt(value_vars=['type1', 'type2'], value_name='type').dropna() # 按类型和原列名分组计数,转宽格式 type_stats = melted.groupby(['type', 'variable']).size().unstack(fill_value=0).reset_index() # 重命名计数列 type_stats = type_stats.rename(columns={'type1': 'type1_count', 'type2': 'type2_count'})
最终会得到包含类型名称、type1计数、type2计数的结构化DataFrame,符合你的需求。
内容的提问来源于stack exchange,提问作者Ivan
相关产品推荐
相关产品推荐

