如何让堆叠条形图按区域总Count降序排列?
问题:按区域总Count降序排列堆叠条形图
数据集
| Month | Region | Count |
|---|---|---|
| April | Bronx | 15 |
| April | Brooklyn | 14 |
| April | Manhattan | 8 |
| April | Nassau County | 2 |
| April | Orange County | 1 |
| April | Out of State | 3 |
| April | Queens | 17 |
| April | Staten Island | 3 |
| April | Unknown | 1 |
| April | Westchester County | 2 |
| July | Bronx | 54 |
| July | Brooklyn | 4 |
| July | Central Valley, NY | 1 |
| July | Clifton, NJ | 1 |
| July | Hackensack, NJ | 1 |
| July | Manhattan | 11 |
| July | Mount Vernon, NY | 1 |
| July | Port Washington, NY | 1 |
| July | Queens | 9 |
| July | Rahway. NJ | 1 |
| July | Staten Island | 1 |
| July | Wappingers Fall, NY | 1 |
| July | Yonker, NY | 3 |
| June | Bronx | 2 |
| June | Brooklyn | 31 |
| June | New Jersey | 2 |
| June | Orange County | 1 |
| June | Queens | 5 |
| March | Bronx | 5 |
| March | Brooklyn | 8 |
| March | Manhattan | 1 |
| March | Nassau County | 4 |
| March | Out of State | 2 |
| March | Queens | 36 |
| March | Suffolk County | 1 |
| March | Unknown | 4 |
| March | Westchester County | 1 |
| May | Bronx | 2 |
| May | Brooklyn | 7 |
| May | Manhattan | 1 |
| May | Nassau County | 2 |
| May | Out of State | 1 |
| May | Queens | 6 |
| May | Staten Island | 17 |
| May | Unknown | 3 |
| May | Westchester County | 1 |
当前实现代码
#chart 1 - attendence by region #making the dataframe shown above s=df.groupby(['Month','Borough']).size().reset_index(name='Count') #s=s.sort_values(['Count'], ascending=False).reset_index() df1=pd.DataFrame(s) #plotting graph sns.set(font_scale = 1.5) sns.set_style("whitegrid", {'font.family':'Century Gothic', 'font.serif':['Century Gothic']}) ax = df1.pivot_table(index='Borough', columns='Month', values='Count', aggfunc='sum', sort=False).plot(kind='barh',grid=False, stacked=True, figsize=(15,10), width=0.9) #bar labels for c in ax.containers: ax.bar_label(c, fmt=lambda x: int(x) if x>0 else '', label_type='center') plt.xlabel('Count') plt.ylabel('Regions') plt.title('<b>Attendence By Region</b>') plt.legend(title='Months',loc='upper right') plt.rcParams["figure.autolayout"] = True
需求
当前图表按区域字母顺序排列,需要改为按各区域**总Count(所有月份之和)**降序排列,此前尝试排序未生效,需解决该问题。
解决方案
问题原因
之前对df1的排序无效,因为pivot_table生成的新表不会继承原DataFrame的排序规则,sort=False仅表示透视表不自动按索引排序,但默认仍会按区域字母顺序排列。正确做法是先计算每个区域的总Count,以此为依据确定区域顺序,再应用到透视表中。
修改后的代码
# chart 1 - attendence by region import pandas as pd import seaborn as sns import matplotlib.pyplot as plt # 生成数据集(假设原df已存在) s = df.groupby(['Month','Borough']).size().reset_index(name='Count') df1 = pd.DataFrame(s) # 计算每个区域的总Count,按降序排序得到区域顺序 region_totals = df1.groupby('Borough')['Count'].sum().sort_values(ascending=False) sorted_regions = region_totals.index.tolist() # 生成透视表并按排序后的区域重新索引 pivot_df = df1.pivot_table( index='Borough', columns='Month', values='Count', aggfunc='sum', sort=False ).reindex(sorted_regions) # 强制使用排序后的区域顺序 # 绘图 sns.set(font_scale=1.5) sns.set_style("whitegrid", {'font.family':'Century Gothic', 'font.serif':['Century Gothic']}) ax = pivot_df.plot(kind='barh', grid=False, stacked=True, figsize=(15,10), width=0.9) # 条形标签 for c in ax.containers: ax.bar_label(c, fmt=lambda x: int(x) if x>0 else '', label_type='center') plt.xlabel('Count') plt.ylabel('Regions') plt.title('<b>Attendence By Region</b>') plt.legend(title='Months', loc='upper right') plt.rcParams["figure.autolayout"] = True plt.show()
关键步骤说明
- 计算区域总Count并排序:通过
groupby('Borough')['Count'].sum()得到每个区域的总数,再用sort_values(ascending=False)降序排列,提取排序后的区域列表。 - 重新索引透视表:对
pivot_table生成的结果使用reindex(sorted_regions),强制按排序后的区域顺序排列,确保绘图时遵循该顺序。 - 保留原绘图逻辑,仅修改数据排序部分,不影响图表样式和标签展示。
内容的提问来源于stack exchange,提问作者JJ.301.
相关产品推荐
相关产品推荐

