如何在同一图表中绘制五个姓名频率DataFrame的折线图?
合并多个姓名频率折线图到同一图表的方法
问题背景
需求:创建图表展示2018-2022年Sophia、Isabella、Olivia、Emma、Mia这五个姓名的频率,已通过以下Pandas代码分别得到对应DataFrame:
Sophia = df_years[(df_years.Childs_First_Name == 'Sophia')] df_Sophia = Sophia.groupby(['Year_of_Birth','Childs_First_Name'])['Count'].sum().reset_index().sort_values(by='Count', ascending=False).head(5) Isabella = df_years[(df_years.Childs_First_Name == 'Isabella')] df_Isabella = Isabella.groupby(['Year_of_Birth','Childs_First_Name'])['Count'].sum().reset_index().sort_values(by='Count', ascending=False).head(5) Olivia = df_years[(df_years.Childs_First_Name == 'Olivia')] df_Olivia = Olivia.groupby(['Year_of_Birth','Childs_First_Name'])['Count'].sum().reset_index().sort_values(by='Count', ascending=False).head(5) Emma = df_years[(df_years.Childs_First_Name == 'Emma')] df_Emma = Emma.groupby(['Year_of_Birth','Childs_First_Name'])['Count'].sum().reset_index().sort_values(by='Count', ascending=False).head(5) Mia = df_years[(df_years.Childs_First_Name == 'Mia')] df_Mia = Mia.groupby(['Year_of_Birth','Childs_First_Name'])['Count'].sum().reset_index().sort_values(by='Count', ascending=False).head(5)当前仅能单独绘制单个DataFrame的折线图(如
df_Mia.plot( 'Year_of_Birth' , 'Count' )),需要将五个DataFrame的折线图合并到同一图表中。
解决方案
方法一:合并DataFrame后统一绘图
这种方法更高效,先将所有小DataFrame合并,再通过透视表整理成适合绘图的格式,最后一次性生成折线图:
import pandas as pd import matplotlib.pyplot as plt # 合并五个姓名的DataFrame combined_df = pd.concat([df_Sophia, df_Isabella, df_Olivia, df_Emma, df_Mia]) # 转换为透视表:行=年份,列=姓名,值=频率 pivot_df = combined_df.pivot( index='Year_of_Birth', columns='Childs_First_Name', values='Count' ).fillna(0) # 填充可能的空值为0 # 绘制折线图 pivot_df.plot(kind='line', figsize=(10, 6)) plt.title('2018-2022年热门姓名频率变化') plt.xlabel('出生年份') plt.ylabel('频率') plt.legend(title='姓名') plt.show()
方法二:在同一坐标轴上逐个绘制
如果不想合并DataFrame,可以先创建一个Matplotlib坐标轴,然后依次在该轴上绘制每个姓名的折线:
import matplotlib.pyplot as plt # 创建画布和坐标轴 fig, ax = plt.subplots(figsize=(10, 6)) # 逐个绘制每个姓名的折线,指定同一个坐标轴ax df_Sophia.plot(x='Year_of_Birth', y='Count', ax=ax, label='Sophia') df_Isabella.plot(x='Year_of_Birth', y='Count', ax=ax, label='Isabella') df_Olivia.plot(x='Year_of_Birth', y='Count', ax=ax, label='Olivia') df_Emma.plot(x='Year_of_Birth', y='Count', ax=ax, label='Emma') df_Mia.plot(x='Year_of_Birth', y='Count', ax=ax, label='Mia') # 添加图表标注 ax.set_title('2018-2022年热门姓名频率变化') ax.set_xlabel('出生年份') ax.set_ylabel('频率') ax.legend(title='姓名') plt.show()
额外优化:简化初始数据处理代码
你当前的代码重复了五次筛选分组逻辑,可以一次性处理所有目标姓名,减少冗余:
target_names = ['Sophia', 'Isabella', 'Olivia', 'Emma', 'Mia'] # 一次性筛选所有目标姓名并分组求和 combined_df = df_years[df_years['Childs_First_Name'].isin(target_names)]\ .groupby(['Year_of_Birth', 'Childs_First_Name'])['Count'].sum().reset_index() # 后续直接用透视表绘图即可 pivot_df = combined_df.pivot(index='Year_of_Birth', columns='Childs_First_Name', values='Count').fillna(0) pivot_df.plot(kind='line', figsize=(10,6)) plt.title('2018-2022年热门姓名频率变化') plt.xlabel('出生年份') plt.ylabel('频率') plt.legend(title='姓名') plt.show()
内容的提问来源于stack exchange,提问作者Matt
相关产品推荐
相关产品推荐

