Python Pandas sort_values排序异常:输出重复行而非预期结果
问题原因
你的代码出现重复行的核心问题是:遍历的是未去重的区域列。country_features['world_region']里每个区域会对应多个国家行,循环时同一个区域会被反复计算均值和数量,最终region_happines里会生成大量重复的区域行,排序后取前5自然都是重复的同一高幸福度区域。
修复方案
方案1:修改原循环(去重区域列表)
只需要在获取区域列表时先去重,确保每个区域只被计算一次:
# 获取去重后的唯一区域列表 regions = country_features['world_region'].unique() happines = [] counts = [] reg = [] for region in regions: hap = country_features.loc[country_features['world_region'] == region, 'happiness_score'].mean() count = len(country_features[country_features['world_region'] == region]) happines.append(hap) counts.append(count) reg.append(region) region_happines = pd.DataFrame({'region':reg, 'happiness_score' : happines, 'country_count':counts}) # 注:mean()返回的已是数值类型,无需额外转格式 sorted_df = region_happines.sort_values(by='happiness_score', ascending=False) sorted_df.head(5)
方案2:用pandas分组聚合(更简洁高效,推荐)
直接利用pandas的groupby完成分组计算,从根源避免重复行问题:
# 一步完成分组、聚合、列重命名 region_happines = country_features.groupby('world_region').agg( happiness_score=('happiness_score', 'mean'), country_count=('world_region', 'count') ).reset_index().rename(columns={'world_region':'region'}) # 排序并取前5 sorted_df = region_happines.sort_values(by='happiness_score', ascending=False) sorted_df.head(5)
内容的提问来源于stack exchange,提问作者ISO
相关产品推荐
相关产品推荐

