如何借助历史年份数据模式开展新会员日新增量预测?
Great question! Your observation that 2017 and 2018 data are highly correlated is a key insight—let’s build on that with more flexible methods that outperform a simple average ratio, while avoiding the flat-line issue with ARIMA and poor Holt-Winter results.
Top Recommended Methods
1. Weighted Year-over-Year Ratio
Instead of taking a single average ratio across all 2018 vs 2017 dates, use weighted averaging to prioritize more recent data. For example:
- Assign higher weights to ratios from the last 3-6 months of 2018 (e.g., 0.7 weight for the final quarter, 0.3 for earlier quarters)
- Use exponential weighted moving average (EWMA) for the ratio sequence, where each day’s ratio gets exponentially less weight as you go back in time. This ensures your forecast reflects the latest trends in the year-over-year relationship, not just a static average.
2. Stratified Day-of-Week Matching
Even if there’s no overall seasonality, day-of-week patterns might be hidden in your average ratio. Here’s how to leverage this:
- Group your data by day of the week (Monday, Tuesday, etc.)
- Calculate a separate 2018/2017 ratio for each group
- For your forecast month, multiply each 2017 day’s value by its corresponding day-of-week ratio. This adjusts for consistent weekly fluctuations that a global average would smooth over.
3. Residual ARIMA on YoY Ratios
Your ratio sequence itself might have subtle autocorrelation that ARIMA can capture (without producing a flat line):
- Compute the daily ratio
r_t = 2018_t / 2017_tfor all overlapping dates - Fit an ARIMA(p,d,q) model to the
r_tsequence (use ACF/PACF plots to pick p and q values) - Forecast the ratios for your target month
- Multiply each forecasted ratio by the corresponding 2017 day’s value to get your final signup predictions. This lets the ratio evolve dynamically instead of being fixed.
4. Regression-Based Forecasting
A simple linear regression can add flexibility beyond a pure ratio:
- Treat 2017 daily signups as the independent variable (
X) and 2018 daily signups as the dependent variable (Y) - Fit the model
Y = a + b*X(whereais an intercept andbis the slope) - For your forecast, plug in the 2017 target month’s values into the equation:
Y_forecast = a + b*X_2017_target. This handles edge cases (like days with zero signups in 2017) where a ratio would be undefined, and accounts for any consistent baseline shift between the two years.
Why These Work Better Than Your Current Approach
All these methods preserve the core strength of your original idea (leveraging high year-over-year correlation) but fix its biggest limitation: a static average ratio can’t account for subtle shifts in the relationship over time or hidden patterns like weekly cycles. They’ll produce more realistic, non-flat forecasts that align with your data’s actual behavior.
内容的提问来源于stack exchange,提问作者Nicholas Humphrey

