如何为含23300条小时级数据的Facebook Prophet模型配置initial、period、horizon及cutoffs参数?
initial, period, horizon and cutoffs for Hourly Data Great question—let’s break this down specifically for your ~23,300 hourly data points, since time granularity makes a big difference in parameter choices.
Configuring initial, period, horizon
These three parameters work together to define how your cross-validation splits historical data, and none require strict adherence to dataset proportions—they should align with your prediction goals and data patterns.
horizon: Prioritize your actual prediction need
This is the length of time you want to forecast into the future, and it doesn’t need to be a fixed percentage of your dataset. For example:
- If your use case requires predicting the next 7 days of hourly data, set
horizon='168 hours'directly. - A good rule of thumb: don’t set
horizonlonger than the period where you trust your model’s predictive power (Prophet’s accuracy drops as the forecast horizon extends). Since you have hourly data, stick to hourly units for clarity.
initial: Give the model enough history to learn patterns
This is the amount of historical data used to train the first model. For hourly data, you need to cover all key periodic patterns (daily, weekly, monthly, potentially annual):
- At minimum, set it to cover 2x your longest expected cycle—for example, if weekly cycles (168 hours) are critical, use
initial='336 hours'(2 weeks) to ensure the model picks up that pattern. - With your ~2.66 years of data (23300/8760 ≈ 2.66), a safer bet is to use 1 year of initial data (
initial='8760 hours'). This gives the model enough context to learn annual trends while leaving plenty of data for cross-validation splits.
period: Balance robustness and computational efficiency
This is the interval between consecutive cutoff points (the end of each training window). It controls how many cross-validation splits you’ll have:
- If you set it too small (e.g., 1 hour), you’ll get thousands of splits and massive computation time.
- If you set it too large, you’ll have too few splits to get a reliable performance estimate.
- For a
horizon='168 hours'(7 days), a good middle ground isperiod='72 hours'(3 days). This creates overlapping validation windows that cover different time points without redundant computation.
cutoffs: Uniform is fine, but custom works too
You don’t need to stick to uniformly distributed cutoffs—they can be fully customized based on your data’s unique characteristics:
- Uniform cutoffs (like the official example) are a safe default: they ensure your validation covers the entire time range of your data, giving you a holistic view of model performance.
- Custom cutoffs are useful if you have specific high-impact time periods (e.g., holidays, sales spikes). For example, if your data has a major annual event in November, add cutoffs in late October to test how well the model predicts the event’s aftermath.
- Just remember: every cutoff must leave at least
horizonhours of data after it (otherwise, there’s no future data to validate against).
Example Code for Your Dataset
Here’s how this might look in practice for a 7-day forecast:
from fbprophet.diagnostics import cross_validation # Assume your trained Prophet model is stored in variable `m` df_cv = cross_validation( m, initial='8760 hours', # 1 year of training data for the first split period='72 hours', # Shift training window forward every 3 days horizon='168 hours' # Forecast 7 days ahead )
For custom cutoffs targeting key dates:
import pandas as pd # Define cutoffs at the start of 2022, mid-2022, and start of 2023 cutoffs = pd.to_datetime([ '2022-01-01 00:00:00', '2022-06-01 00:00:00', '2023-01-01 00:00:00' ]) df_cv2 = cross_validation(m, cutoffs=cutoffs, horizon='168 hours')
内容的提问来源于stack exchange,提问作者Tobitor

