使用Python Plotly创建多分类图表时子类别顺序混乱求助
问题
我有如下Excel格式数据:包含company、month-year、#people got interviewed、# people employed字段。尝试用Python的Plotly库创建以company为一级分类、month-year为二级分类的多分类柱状图时,Y、Z公司的二级分类顺序出现混乱。我已将日期字符串转为datetime对象排序后转回字符串,打印顺序正确,但图表仍混乱,求解释原因。
代码如下:
import pandas as pd from helper_functions import get_df import plotly.graph_objects as go from datetime import datetime def multicat_chart(infile=None, sheet_name=None, chart_type = None, chart_title = None): #chart type must be given df=pd.read_excel(infile,sheet_name) df = df.fillna(method='ffill') cat = df.columns[0] sub_cat = df.columns[1] cols = df.columns[2:] fig = go.Figure() cats = [] sub_cats = [] for c in df[cat].unique(): new_df = df.loc[df[cat] == c] scats = new_df[sub_cat] scats = scats.apply(lambda date: datetime.strptime(date, "%b-%Y")) scats = list(scats) scats.sort() scats = [datetime.strftime(element, '%b-%y') for element in scats] scats = [str(element) for element in scats] for sc in scats: cats.append(str(c)) sub_cats.append(str(sc)) print(c) for i in scats: print(i) fig.add_trace( go.Bar(x = [cats,sub_cats],y = df[cols[0]], name="# people got interviewed" )) fig.add_trace( go.Bar(x = [cats,sub_cats],y = df[cols[1]], name="# people employed" )) fig.update_layout(width = 1000, height = 1000) return fig fig = multicat_chart(infile = 'data_for_test.xlsx', sheet_name = 'data', chart_type = 'bar') fig.show()
解答
问题核心是你只单独排序了日期字符串,但没有同步调整对应y轴数据的顺序,导致图表中x轴的分类顺序和y轴数值错位,最终显示混乱。
具体细节:
- 循环每个公司时,你对该公司的
month-year做了排序,生成了有序的cats和sub_cats列表,但传给图表的df[cols[0]]和df[cols[1]]依然是原始数据的顺序,和排序后的分类没有对应关系。 - 打印的是排序后的日期,但实际图表使用的y值还是原始数据行的顺序,两者不匹配,所以二级分类显示混乱。
解决方法:
需要对每个公司的子数据框整体按日期排序,再提取对应分类和数值:
import pandas as pd from helper_functions import get_df import plotly.graph_objects as go from datetime import datetime def multicat_chart(infile=None, sheet_name=None, chart_type = None, chart_title = None): df=pd.read_excel(infile,sheet_name) df = df.fillna(method='ffill') cat = df.columns[0] sub_cat = df.columns[1] cols = df.columns[2:] fig = go.Figure() cats = [] sub_cats = [] interviewed_vals = [] employed_vals = [] for c in df[cat].unique(): new_df = df.loc[df[cat] == c].copy() # 将month-year转为datetime并对整个子数据框排序 new_df['sorted_date'] = pd.to_datetime(new_df[sub_cat], format="%b-%Y") new_df = new_df.sort_values('sorted_date') # 提取排序后的分类和对应数值 sorted_subcats = new_df[sub_cat].apply(lambda x: datetime.strptime(x, "%b-%Y").strftime('%b-%y')).tolist() cats.extend([str(c)] * len(sorted_subcats)) sub_cats.extend(sorted_subcats) interviewed_vals.extend(new_df[cols[0]].tolist()) employed_vals.extend(new_df[cols[1]].tolist()) # 打印验证 print(c) for i in sorted_subcats: print(i) fig.add_trace( go.Bar(x = [cats,sub_cats],y = interviewed_vals, name="# people got interviewed" )) fig.add_trace( go.Bar(x = [cats,sub_cats],y = employed_vals, name="# people employed" )) fig.update_layout(width = 1000, height = 1000) return fig fig = multicat_chart(infile = 'data_for_test.xlsx', sheet_name = 'data', chart_type = 'bar') fig.show()
修改后,x轴的分类顺序和y轴数值一一对应,图表就能正确显示有序的二级分类。
内容的提问来源于stack exchange,提问作者Tigran Mkrtchyan
相关产品推荐
相关产品推荐

