多部门季度数据预测:Plotly多Y轴折线图报错修复
问题描述
需要对Fin、Legal、Leadership、Overall等部门的季度数据进行时间序列预测,直到2023年第四季度。尝试用Plotly绘制多变量折线图时触发ValueError,错误信息如下:
ValueError: All arguments should have the same length. The length of argument y is 4, whereas the length of previously-processed arguments ['Time Period'] is 6
示例数据
| Time Period HR | Fin | Legal | Leadership | Overall |
|---|---|---|---|---|
| 2021Q2 | 42 | 36 | 66 | 53 |
| 2021Q3 | 52 | 43 | 64 | 67 |
| 2021Q4 | 65 | 47 | 71 | 73 |
| 2022Q1 | 68 | 50 | 75 | 74 |
| 2022Q2 | 72 | 57 | 77 | 81 |
| 2022Q3 | 79 | 62 | 75 | 78 |
报错代码
import pandas as pd from datetime import date, timedelta import datetime import matplotlib.pyplot as plt plt.style.use('fivethirtyeight') from statsmodels.tsa.seasonal import seasonal_decompose from statsmodels.graphics.tsaplots import plot_pacf from statsmodels.tsa.arima_model import ARIMA import statsmodels.api as sm import warnings import plotly.graph_objects as go from plotly.subplots import make_subplots inputfilepath = 'C:/Documents/Forecast/Input/Forecast Data csv.csv' df = pd.read_csv(inputfilepath) print(df) import plotly.express as px figure = px.line(df, x="Time Period", y=("Fin","Legal","Leadership","Overall"), title='Quarterly scores') figure.show()
问题分析与修复
错误原因
报错核心是DataFrame列名不匹配:示例数据第一列的实际列名为Time Period HR(疑似输入笔误),但代码中指定的x轴列名为Time Period,导致Plotly无法找到对应列,进而误判数据长度不匹配。
修复后的绘图代码
先修正列名,再绘制多变量折线图:
import pandas as pd import plotly.express as px # 读取数据 inputfilepath = 'C:/Documents/Forecast/Input/Forecast Data csv.csv' df = pd.read_csv(inputfilepath) # 修正列名(匹配代码中使用的x轴字段) df.rename(columns={"Time Period HR": "Time Period"}, inplace=True) # 绘制多Y轴折线图 figure = px.line(df, x="Time Period", y=["Fin", "Legal", "Leadership", "Overall"], title='各部门季度得分趋势') # 优化轴标签显示 figure.update_layout(xaxis_title="时间周期", yaxis_title="得分") figure.show()
扩展:实现2023Q4前的时间序列预测
针对季度季节性数据,使用SARIMA模型完成预测,完整代码如下:
import pandas as pd import warnings import plotly.graph_objects as go from statsmodels.tsa.statespace.sarimax import SARIMAX warnings.filterwarnings("ignore") # 1. 数据预处理 inputfilepath = 'C:/Documents/Forecast/Input/Forecast Data csv.csv' df = pd.read_csv(inputfilepath) df.rename(columns={"Time Period HR": "Time Period"}, inplace=True) # 将季度字符串转为datetime格式,设置时间索引 df['Time Period'] = pd.to_datetime(df['Time Period'].str.replace('Q', '-') + '-1', format='%Y-%m-%d') df.set_index('Time Period', inplace=True) df = df.asfreq('Q') # 明确季度频率 # 2. 定义单部门预测函数(适配季度季节性) def forecast_dept(data_col, steps=5): # SARIMA(1,1,1)(1,1,1,4):4代表季度周期 model = SARIMAX(data_col, order=(1,1,1), seasonal_order=(1,1,1,4)) result = model.fit() # 生成预测值与置信区间 forecast = result.get_forecast(steps=steps) return forecast.predicted_mean, forecast.conf_int() # 3. 预测各部门至2023Q4(现有数据到2022Q3,需预测5个季度) departments = ["Fin", "Legal", "Leadership", "Overall"] forecast_data = {} for dept in departments: forecast_data[dept] = forecast_dept(df[dept], steps=5) # 4. 绘制含预测值的多折线图 fig = go.Figure() # 添加历史数据轨迹 for dept in departments: fig.add_trace(go.Scatter(x=df.index, y=df[dept], mode='lines+markers', name=f'{dept} 历史数据')) # 添加预测数据及置信区间 for dept in departments: pred_vals, conf_int = forecast_data[dept] fig.add_trace(go.Scatter(x=pred_vals.index, y=pred_vals, mode='lines+markers', name=f'{dept} 预测数据', line=dict(dash='dash'))) # 填充置信区间 fig.add_trace(go.Scatter( x=pred_vals.index.append(pred_vals.index[::-1]), y=conf_int.iloc[:,0].append(conf_int.iloc[:,1][::-1]), fill='toself', fillcolor='rgba(0,0,0,0.1)', line=dict(color='rgba(255,255,255,0)'), name=f'{dept} 置信区间' )) # 优化图表布局 fig.update_layout(title='各部门季度得分及预测(至2023Q4)', xaxis_title='时间周期', yaxis_title='得分', legend=dict(orientation='h', yanchor='bottom', y=1.02, xanchor='right', x=1)) fig.show()
内容的提问来源于stack exchange,提问作者Scythor
相关产品推荐
相关产品推荐

