时间维度重要时,随机拆分时间序列是否适用于NARX等神经网络训练?
Great question—this is a critical, often misunderstood pain point when working with sequence-dependent models like NARX, so let’s unpack it thoroughly.
Short Answer
No, random splitting is almost never appropriate when time dimension is important for NARX or any time-series-focused neural network. The claim that "differences from resampling are larger than those from network conditions" is absolutely valid, and here’s why:
Why Random Splitting Fails for Time-Series Data
Time-series data has inherent temporal dependency—each data point is correlated with past (and sometimes future) points. Randomly shuffling and splitting breaks this structure in two harmful ways:
- Future data leakage: You’ll end up with training samples that include data points from later time periods than some validation/test samples. NARX networks learn to map historical inputs to outputs, so this leakage lets the model "cheat" by implicitly using future information, leading to inflated performance metrics that don’t reflect real-world generalization.
- Broken sequence patterns: NARX relies on learning sequential relationships (e.g., how yesterday’s values influence today’s output). Random splitting destroys these patterns, forcing the model to learn noise instead of meaningful temporal dynamics.
The Resampling vs. Network Conditions Claim
That quote hits the nail on the head. When you use random splitting, the variance in model performance across different random seeds (i.e., different splits) will dwarf any differences caused by tuning network hyperparameters (like number of hidden units, learning rate) or architecture tweaks. This is because:
- A random split might accidentally include a major seasonal spike or anomaly in the training set, making the model look great—while another split excludes that spike, leading to terrible performance.
- The core issue isn’t the network itself; it’s that the split doesn’t reflect how the model will be used in production (where you’ll always predict future data from past observations).
When Might Random Splitting Be Okay?
The only scenario where random splitting could work is if your time-series is strictly stationary—no trend, no seasonality, and no significant temporal correlation between consecutive points. But let’s be real: if your data fits that description, you probably don’t need a NARX network in the first place. NARX is designed for data with strong sequential dependencies, so this edge case is rare.
Recommended Splitting Approaches for NARX
Stick to time-ordered splits to keep temporal integrity intact:
- Simple time-based split: Use the first 80% of your chronological data for training, next 10% for validation (to tune hyperparameters), and final 10% for testing (to evaluate final performance).
- Time-series cross-validation: For smaller datasets, use a rolling window approach. For example, train on weeks 1–4, validate on week 5; then train on weeks 1–5, validate on week 6; repeat until you cover the dataset. This gives you a more robust estimate of model performance.
Final Takeaway
Random splitting undermines the core purpose of NARX networks—learning from sequential data. The resampling method has a far bigger impact on your model’s real-world utility than any network configuration choices. Always prioritize time-ordered splits when time dimension is critical.
内容的提问来源于stack exchange,提问作者Rom

