You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PyTorch-Forecasting中TimeSeriesDataSet的使用疑问

时间序列训练集问题解答

背景信息

我拥有913000行销售数据,涵盖10家门店、50个商品2013-01-01至2017-12-31的销售记录(因闰年产生该行数)。创建训练集的代码如下:

training = TimeSeriesDataSet(
    train_df[train_df.apply(lambda x:x['time_idx']<=training_cutoff,axis=1)],    
    time_idx = "time_idx",
    target = "sales",
    group_ids = ["store","item"], # list of column names identifying a time series
    max_encoder_length = max_encoder_length,
    max_prediction_length = max_prediction_length,
    static_categoricals = ["store","item"],
    # Categorical variables that do not change over time (e.g. product length)
    time_varying_unknown_reals = ["sales"],
    
)

问题1解答

你的理解不完全准确。TimeSeriesDataSet的输入是你筛选出的time_idx<=training_cutoff的原始数据,但最终生成的有效训练样本数还要受每个时间序列组(store+item)的长度限制:

  • 每个组至少需要max_encoder_length个连续历史数据点,才能生成一个可用于训练的样本
  • 即使原始数据在training_cutoff范围内,组内数据长度不足max_encoder_length的部分(比如每个组的前max_encoder_length-1行)会被自动过滤,无法转化为有效训练样本

问题2解答

差值9500的核心原因是单组时间序列的有效样本过滤规则,具体分析:

  1. 你的计算(913000-10000-30000=873000)假设所有符合time_idx<=training_cutoff的数据都能转化为有效样本,但实际并非如此
  2. 你共有500个独立时间序列组(10家门店×50个商品),每个组需要满足:只有当组内已有至少max_encoder_length个数据点时,才能从该组生成训练样本。那些组内数据长度不足max_encoder_length的行(比如每个组的前N行,N=max_encoder_length-1)会被剔除
  3. 结合数值来看,你实际总剔除行数是913000-863500=49500,比预期的40000多了9500,这说明:
    • 要么是部分组的起始数据晚于2013-01-01,导致组内有效数据长度更短
    • 要么是部分组存在数据缺失(中间断档),使得连续可用的历史数据不足max_encoder_length
  4. 验证方法:统计每个store+item组的有效数据行数,计算每个组可生成的样本数(组内行数 - max_encoder_length +1,若结果为正),求和后即可得到实际训练集的有效样本数

内容的提问来源于stack exchange,提问作者sungjoonlee

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.10 01:15:37