PyTorch-Forecasting中TimeSeriesDataSet的使用疑问
时间序列训练集问题解答
背景信息
我拥有913000行销售数据,涵盖10家门店、50个商品2013-01-01至2017-12-31的销售记录(因闰年产生该行数)。创建训练集的代码如下:
training = TimeSeriesDataSet( train_df[train_df.apply(lambda x:x['time_idx']<=training_cutoff,axis=1)], time_idx = "time_idx", target = "sales", group_ids = ["store","item"], # list of column names identifying a time series max_encoder_length = max_encoder_length, max_prediction_length = max_prediction_length, static_categoricals = ["store","item"], # Categorical variables that do not change over time (e.g. product length) time_varying_unknown_reals = ["sales"], )
问题1解答
你的理解不完全准确。TimeSeriesDataSet的输入是你筛选出的time_idx<=training_cutoff的原始数据,但最终生成的有效训练样本数还要受每个时间序列组(store+item)的长度限制:
- 每个组至少需要
max_encoder_length个连续历史数据点,才能生成一个可用于训练的样本 - 即使原始数据在
training_cutoff范围内,组内数据长度不足max_encoder_length的部分(比如每个组的前max_encoder_length-1行)会被自动过滤,无法转化为有效训练样本
问题2解答
差值9500的核心原因是单组时间序列的有效样本过滤规则,具体分析:
- 你的计算(913000-10000-30000=873000)假设所有符合
time_idx<=training_cutoff的数据都能转化为有效样本,但实际并非如此 - 你共有500个独立时间序列组(10家门店×50个商品),每个组需要满足:只有当组内已有至少
max_encoder_length个数据点时,才能从该组生成训练样本。那些组内数据长度不足max_encoder_length的行(比如每个组的前N行,N=max_encoder_length-1)会被剔除 - 结合数值来看,你实际总剔除行数是913000-863500=49500,比预期的40000多了9500,这说明:
- 要么是部分组的起始数据晚于2013-01-01,导致组内有效数据长度更短
- 要么是部分组存在数据缺失(中间断档),使得连续可用的历史数据不足
max_encoder_length
- 验证方法:统计每个
store+item组的有效数据行数,计算每个组可生成的样本数(组内行数 - max_encoder_length +1,若结果为正),求和后即可得到实际训练集的有效样本数
内容的提问来源于stack exchange,提问作者sungjoonlee
相关产品推荐
相关产品推荐

