You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

能否使用mlr包基于面板研究数据结合时间维度完成预测任务?

纵向面板数据的mlr包机器学习预测问题

我计划开展一项监督机器学习项目,使用纵向研究(面板研究)的数据,目标是利用2004年和2009年的预测变量预测2014年的结果。目前已完成初步数据预处理,数据框的简化结构如下:

data_ml <- structure(
  list(
    ID = c(
      201,
      203,
      602,
      901,
      1202,
      1501,
      1601,
      1602,
      1603,
      201,
      203,
      602,
      901,
      1202,
      1501,
      1601,
      1602,
      1603,
      201,
      203,
      602,
      901,
      1202,
      1501,
      1601,
      1602,
      1603
    ),
    Studyyear = c(
      2004,
      2004,
      2004,
      2004,
      2004,
      2004,
      2004,
      2004,
      2004,
      2009,
      2009,
      2009,
      2009,
      2009,
      2009,
      2009,
      2009,
      2009,
      2014,
      2014,
      2014,
      2014,
      2014,
      2014,
      2014,
      2014,
      2014
    ),
    Gender = c(2, 1, 2, 2, 2, 1, 1, 2, 1,
               2, 1, 2, 2, 2, 1, 1, 2, 1, 2, 1, 2, 2, 2, 1, 1, 2, 1),
    Predictor1 = c(6,
                   5, 4, 6, 4, 6, 4, 3, 3, 6, 5, 4, 6, 4, 6, 4, 3, 3, 6, 5, 4, 6,
                   4, 6, 4, 3, 3),
    Predictor2 = c(2, 2, 1, 1, 2, 2, 1, 2, 2, 2,
                   2, 1, 1, 2, 2, 1, 2, 2, 2, 2, 1, 1, 2, 2, 1, 2, 2),
    Predictor3 = c(0,
                   6, 1, 6, 0, 0, 4, 2, 3, 0, 6, 1, 6, 0, 0, 4, 1, 1, 1, 6, 1, 6,
                   0, 0, 4, 1, 1),
    Outcome1 = c(0, 1, 1, 0, 0, 0, 0, 0, 1, 0, 1,
                 1, 0, 0, 1, 1, 1, 0, 1, 1, 0, 0, 1, 0, 0, 1, 1),
    Outcome2 = c(0,
                 0, 1, 0, 0, 1, 0, 0, 1, 0, 0, 1, 0, 0, 0, 1, 1, 0, 1, 1, 1, 0,
                 1, 0, 1, 1, 0)
  ),
  class = c("tbl_df", "tbl", "data.frame"),
  row.names = c(NA,-27L)
)

此前我的预测项目未涉及时间维度,曾按如下方式使用mlr包创建任务并进行预测:

library(mlr)
task <- makeClassifTask(data = data_ml, target = 'Outcome1', positive = '1')
measures = list(acc, auc, tpr, tnr, f1)
resampling_MC <- makeResampleDesc(method = 'Subsample', iters = 500) 
learner_logreg <- makeLearner('classif.logreg', predict.type = 'prob')
benchmark_MC <- benchmark(learners = learner_logreg, tasks = task, resamplings = resampling_MC, measures = measures)

请问针对上述包含Studyyear时间维度的数据框,仍可使用mlr包完成预测任务吗?


回答

完全可以用mlr包完成这个纵向面板数据的预测任务,但需要针对时间维度调整数据结构和重采样策略,避免数据泄露问题:

  1. 重构数据结构
    当前数据是长格式(每个ID对应3行不同年份数据),需要转换为宽格式:每个ID一行,包含2004和2009年的预测变量(比如Predictor1_2004、Predictor1_2009),以及2014年的结果变量(Outcome1_2014、Outcome2_2014)。

    示例转换代码:

    library(tidyr)
    data_wide <- data_ml %>%
      pivot_wider(
        id_cols = ID,
        names_from = Studyyear,
        values_from = c(Predictor1, Predictor2, Predictor3, Outcome1, Outcome2),
        names_sep = "_"
      ) %>%
      # 保留用于预测的列和2014年的结果
      select(ID, Gender, starts_with("Predictor"), ends_with("_2014"))
    
  2. 调整重采样策略
    不能再用随机Subsample重采样,否则会导致训练集包含未来年份(2014)的数据,造成数据泄露。可以选择:

    • 按ID分组重采样:确保同一ID的所有数据要么在训练集要么在测试集,避免用同一个体的未来数据训练
    • 时间序列分割:如果有更多批次时间点,用时间顺序划分训练/测试集

    示例分组重采样代码:

    # 创建按ID分组的重采样描述
    resampling_group <- makeResampleDesc("Subsample", iters = 500, group = "ID")
    
  3. 创建分类任务
    基于转换后的宽格式数据创建任务,指定目标变量为2014年的结果:

    task <- makeClassifTask(data = data_wide, target = 'Outcome1_2014', positive = '1')
    
  4. 后续流程与之前一致
    选择学习器、设置评估指标、运行基准测试的步骤和你之前的代码逻辑完全兼容,只需要替换任务和重采样描述即可。


内容的提问来源于stack exchange,提问作者Mangus

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.17 04:45:35