You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python绘制人口金字塔:解决百分比标注重复问题

解决人口金字塔百分比标注重复问题

问题原因

标注重复是因为两次调用sns.barplot后,male_barplot.patches和female_barplot.patches会包含所有年龄组的男性、女性条形(两个图层的条形被合并到同一个patches列表),遍历的时候会重复处理每个条形。而你之前用iterrows报错,大概率是没正确对应分类y轴的坐标位置(分类轴的坐标是从0开始的整数索引,不是年龄组文本)。

解决方案

直接在DataFrame中计算好百分比,基于数据行的索引定位标注位置,避开遍历patches的坑:

步骤1:提前计算百分比

在创建age_p数据框后,添加百分比计算:

population_total = cleandf.shape[0]
# 计算各年龄组男/女性占总人口的百分比
age_p['Male_Pct'] = (age_p['Male'].abs() / population_total) * 100
age_p['Female_Pct'] = (age_p['Female'] / population_total) * 100

步骤2:替换标注逻辑

删掉原来遍历patches的两个循环,换成基于iterrows的标注(此时能正确对应y轴位置):

# 标注男性百分比
for idx, row in age_p.iterrows():
    if row['Male'] != 0:
        x = row['Male'] - 10  # 向左偏移,避免和条形重叠
        y = idx  # 分类y轴的坐标是0、1、2...对应每个年龄组
        male_barplot.annotate(f'{row["Male_Pct"]:.1f}%', (x, y), ha='center', va='center')

# 标注女性百分比
for idx, row in age_p.iterrows():
    if row['Female'] != 0:
        x = row['Female'] + 10  # 向右偏移
        y = idx
        female_barplot.annotate(f'{row["Female_Pct"]:.1f}%', (x, y), ha='center', va='center')

完整修正代码

import pandas as pd
import seaborn as sns
import matplotlib.pyplot as plt

# 假设cleandf为已清洗的数据集
# 统计各年龄组男女人数
male_age_count = cleandf[cleandf['Gender'] == 'Male'].groupby('Age Class')['Age Class'].count()
female_age_count = cleandf[cleandf['Gender'] == 'Female'].groupby('Age Class')['Age Class'].count()

# 排序年龄组
ageclass = cleandf['Age Class'].unique()
ageclass = sorted(ageclass, key=lambda x: 200 if x == '100 +' else int(x.split('-')[0]), reverse=True)

# 构建绘图用数据框
age_p = pd.DataFrame({
    'Age': ageclass,
    'Male': -male_age_count.reindex(ageclass),
    'Female': female_age_count.reindex(ageclass)
})

# 计算百分比
population_total = cleandf.shape[0]
age_p['Male_Pct'] = (age_p['Male'].abs() / population_total) * 100
age_p['Female_Pct'] = (age_p['Female'] / population_total) * 100

# 绘制金字塔
plt.figure(figsize=(10, 6))
male_barplot = sns.barplot(x='Male', y='Age', data=age_p, color='mediumblue', label='Male')
female_barplot = sns.barplot(x='Female', y='Age', data=age_p, color='darkorange', label='Female')

# 添加百分比标注
for idx, row in age_p.iterrows():
    if row['Male'] != 0:
        male_barplot.annotate(f'{row["Male_Pct"]:.1f}%', 
                             (row['Male'] - 10, idx), 
                             ha='center', va='center')
for idx, row in age_p.iterrows():
    if row['Female'] != 0:
        female_barplot.annotate(f'{row["Female_Pct"]:.1f}%', 
                               (row['Female'] + 10, idx), 
                               ha='center', va='center')

# 其他绘图设置
plt.text(-400, 5, 'Male', fontsize=15, fontweight='bold')
plt.text(300, 5, 'Female', fontsize=15, fontweight='bold')
plt.legend(loc='best')
plt.xticks(range(-600, 700, 100))
plt.title('Population/Age Pyramid', fontsize=20, fontweight='bold')
plt.xlabel('Population', fontsize=15, fontweight='bold')
plt.ylabel('Age Range', fontsize=15, fontweight='bold')

plt.show()

说明

  • 直接在DataFrame中计算百分比,避免重复计算,逻辑更清晰。
  • 用iterrows遍历数据时,idx对应分类y轴的整数坐标,精准匹配每个年龄组的条形位置,解决之前iterrows报错的问题。
  • 不再依赖patches列表,彻底避免了两次绘图导致的标注重复问题。

内容的提问来源于stack exchange,提问作者leakie

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.21 16:30:11