Python绘制人口金字塔:解决百分比标注重复问题
解决人口金字塔百分比标注重复问题
问题原因
标注重复是因为两次调用sns.barplot后,male_barplot.patches和female_barplot.patches会包含所有年龄组的男性、女性条形(两个图层的条形被合并到同一个patches列表),遍历的时候会重复处理每个条形。而你之前用iterrows报错,大概率是没正确对应分类y轴的坐标位置(分类轴的坐标是从0开始的整数索引,不是年龄组文本)。
解决方案
直接在DataFrame中计算好百分比,基于数据行的索引定位标注位置,避开遍历patches的坑:
步骤1:提前计算百分比
在创建age_p数据框后,添加百分比计算:
population_total = cleandf.shape[0] # 计算各年龄组男/女性占总人口的百分比 age_p['Male_Pct'] = (age_p['Male'].abs() / population_total) * 100 age_p['Female_Pct'] = (age_p['Female'] / population_total) * 100
步骤2:替换标注逻辑
删掉原来遍历patches的两个循环,换成基于iterrows的标注(此时能正确对应y轴位置):
# 标注男性百分比 for idx, row in age_p.iterrows(): if row['Male'] != 0: x = row['Male'] - 10 # 向左偏移,避免和条形重叠 y = idx # 分类y轴的坐标是0、1、2...对应每个年龄组 male_barplot.annotate(f'{row["Male_Pct"]:.1f}%', (x, y), ha='center', va='center') # 标注女性百分比 for idx, row in age_p.iterrows(): if row['Female'] != 0: x = row['Female'] + 10 # 向右偏移 y = idx female_barplot.annotate(f'{row["Female_Pct"]:.1f}%', (x, y), ha='center', va='center')
完整修正代码
import pandas as pd import seaborn as sns import matplotlib.pyplot as plt # 假设cleandf为已清洗的数据集 # 统计各年龄组男女人数 male_age_count = cleandf[cleandf['Gender'] == 'Male'].groupby('Age Class')['Age Class'].count() female_age_count = cleandf[cleandf['Gender'] == 'Female'].groupby('Age Class')['Age Class'].count() # 排序年龄组 ageclass = cleandf['Age Class'].unique() ageclass = sorted(ageclass, key=lambda x: 200 if x == '100 +' else int(x.split('-')[0]), reverse=True) # 构建绘图用数据框 age_p = pd.DataFrame({ 'Age': ageclass, 'Male': -male_age_count.reindex(ageclass), 'Female': female_age_count.reindex(ageclass) }) # 计算百分比 population_total = cleandf.shape[0] age_p['Male_Pct'] = (age_p['Male'].abs() / population_total) * 100 age_p['Female_Pct'] = (age_p['Female'] / population_total) * 100 # 绘制金字塔 plt.figure(figsize=(10, 6)) male_barplot = sns.barplot(x='Male', y='Age', data=age_p, color='mediumblue', label='Male') female_barplot = sns.barplot(x='Female', y='Age', data=age_p, color='darkorange', label='Female') # 添加百分比标注 for idx, row in age_p.iterrows(): if row['Male'] != 0: male_barplot.annotate(f'{row["Male_Pct"]:.1f}%', (row['Male'] - 10, idx), ha='center', va='center') for idx, row in age_p.iterrows(): if row['Female'] != 0: female_barplot.annotate(f'{row["Female_Pct"]:.1f}%', (row['Female'] + 10, idx), ha='center', va='center') # 其他绘图设置 plt.text(-400, 5, 'Male', fontsize=15, fontweight='bold') plt.text(300, 5, 'Female', fontsize=15, fontweight='bold') plt.legend(loc='best') plt.xticks(range(-600, 700, 100)) plt.title('Population/Age Pyramid', fontsize=20, fontweight='bold') plt.xlabel('Population', fontsize=15, fontweight='bold') plt.ylabel('Age Range', fontsize=15, fontweight='bold') plt.show()
说明
- 直接在DataFrame中计算百分比,避免重复计算,逻辑更清晰。
- 用
iterrows遍历数据时,idx对应分类y轴的整数坐标,精准匹配每个年龄组的条形位置,解决之前iterrows报错的问题。 - 不再依赖
patches列表,彻底避免了两次绘图导致的标注重复问题。
内容的提问来源于stack exchange,提问作者leakie
相关产品推荐
相关产品推荐

