基于ARIMA的136年月度降雨量单变量时间序列预测问题咨询
Hey there! Let's break down and solve the two key problems you're facing with your 136-year monthly rainfall time series prediction using ARIMA:
1. Fixing Negative Forecast Values
ARIMA is a linear model, which means it doesn’t inherently respect the non-negative constraint of rainfall. It extrapolates trends linearly, so if your training data has periods of low rainfall, the model might overshoot into negative territory when forecasting. Here are practical fixes:
- Truncate negative predictions: The simplest quick fix is to set any forecast value < 0 to 0. This works well if negative values are rare and you just need plausible output.
- Log transformation (with adjustment for zeros): Since you have zero values, use
log(rainfall + 1)to transform the data before fitting ARIMA. After forecasting, reverse the transformation withexp(prediction) - 1to get back to the original scale. This pushes predictions firmly into positive territory. - Switch to non-negative time series models: Consider models designed specifically for count/non-negative data, like:
- Zero-Inflated SARIMA (ZISARIMA), which explicitly models the zero-value segments and positive rainfall amounts separately
- Generalized Linear Models (GLMs) with ARIMA errors (e.g., Poisson or negative binomial GLMs), which account for non-negative, intermittent data structures
2. Addressing Zero-to-High Value Jumps & Poor Fit
Your intuition is spot-on—these sudden jumps violate ARIMA’s core assumption that observations depend on smooth, correlated historical values. Rainfall is a classic zero-inflated intermittent time series, where long stretches of zeros (dry seasons) are followed by abrupt high values (monsoons). Here’s how to tackle this:
- Model seasonality explicitly: For monthly data, set the seasonal period in SARIMA to match your region’s monsoon cycle (e.g., 6 months if monsoons run June–November, or 12 for full annual cycles). This helps the model learn when high rainfall is likely to occur, even after extended dry spells.
- Split the data into sub-series: Separate dry season (mostly zeros) and wet season (high values) data. Model each segment independently:
- For dry seasons, predict the probability of zero rainfall
- For wet seasons, use ARIMA or a non-linear model to forecast positive rainfall amounts
- Try non-linear models: Machine learning models like XGBoost, Random Forest, or LSTMs are far better at capturing the non-linear jumps between zero and high values. They don’t rely on strict linear autocorrelation assumptions like ARIMA.
- Why annual/quarterly data didn’t work: Annual data aggregates monthly values, which can mask the intermittent pattern or reduce variability to the point where ARIMA can’t learn meaningful trends. Quarterly data breaks your monsoon period, so the model loses the critical seasonal trigger for high rainfall.
内容的提问来源于stack exchange,提问作者Ankita Fernandes

