使用PyCaret做时间序列预测时compare_models函数卡顿问题排查
问题:PyCaret时间序列预测中compare_models卡在20%无法运行
在Google Colab中使用PyCaret 3.1.0工具,基于汽车经销商零件收入数据集进行时间序列预测时,调用compare_models函数对比模型以寻找最优模型,代码卡在20%进度无法继续(截图显示进度停留在20%)。相关代码如下:
# Only enable critical logging (Optional) import os os.environ["PYCARET_CUSTOM_LOGGING_LEVEL"] = "CRITICAL" def what_is_installed(): from pycaret import show_versions show_versions() try: what_is_installed() except ModuleNotFoundError: !pip install pycaret what_is_installed() import pandas as pd import numpy as np import pycaret pycaret.__version__ # 3.1.0 df = pd.read_csv('parts_revenue.csv', delimiter=';') from pycaret.utils.time_series import clean_time_index cleaned = clean_time_index(data=df, index_col='Posting Date', freq='D') # Verify the resulting DataFrame print(cleaned.head(n=50)) # parts['MA12'] = parts['Parts Revenue'].rolling(12).mean() # import plotly.express as px # fig = px.line(parts, x="Posting Date", y=["Parts Revenue", # "MA12"], template = 'plotly_dark') # fig.show() import time import numpy as np from pycaret.time_series import * # We want to forecast the next 12 days of data and we will use 3 # fold cross-validation to test the models. fh = 12 # or alternately fh = np.arange(1,13) fold = 3 # Global Figure Settings for notebook ---- # Depending on whether you are using jupyter notebook, jupyter lab, # Google Colab, you may have to set the renderer appropriately # NOTE: Setting to a static renderer here so that the notebook # saved size is reduced. fig_kwargs = { # "renderer": "notebook", "renderer": "png", "width": 1000, "height": 600, } """## EDA""" eda = TSForecastingExperiment() eda.setup(cleaned, fh=fh, numeric_imputation_target = 0, fig_kwargs=fig_kwargs ) eda.plot_model() eda.plot_model(plot="diagnostics", fig_kwargs={"height": 800, "width": 1000} ) eda.plot_model( plot="diff", data_kwargs={"lags_list": [[1], [1, 7]], "acf": True, "pacf": True, "periodogram": True}, fig_kwargs={"height": 800, "width": 1500} ) """## Modeling""" exp = TSForecastingExperiment() exp.setup(data = cleaned, fh=fh, numeric_imputation_target = 0.0, fig_kwargs=fig_kwargs, seasonal_period = 5 ) # compare baseline models best = exp_ts.compare_models(errors = 'raise') # CODE HANGS HERE! # plot forecast for 36 months in future plot_model(best, plot = 'forecast', data_kwargs = {'fh' : 24} )
可能的原因及解决办法
- 变量名错误:建模部分初始化了
exp = TSForecastingExperiment(),但调用对比模型时使用了未定义的exp_ts,应改为exp.compare_models,这是直接导致代码异常的核心原因。 - 重复实验初始化:先后初始化
eda和exp两个实验对象并分别执行setup,可能造成资源占用或环境冲突,建议复用同一个实验对象,或在新实验前清理环境。 - 模型计算负载过高:
compare_models默认会训练大量时间序列模型,Colab免费算力下,复杂模型(如深度学习模型)训练耗时极长,看似卡住实则在后台运行。可通过include参数指定轻量模型,比如:best = exp.compare_models(include=['arima', 'ets', 'fbprophet'], errors='raise') - 参数与数据不匹配:设置
seasonal_period=5但数据是日度频率,若实际业务周期为周(7天)或月,参数不匹配会导致模型训练异常;另外numeric_imputation_target=0的填充逻辑是否合理?若为缺失值,插值填充可能更符合业务逻辑,不当的预处理也会影响模型训练。
内容的提问来源于stack exchange,提问作者AndCh
相关产品推荐
相关产品推荐

