You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何构建含NA的ID级跨时间点暴露-结局配对数据集?

解决方案:构建含time_1暴露与time_3结局的宽格式数据集

需求说明

将长格式数据集转换为每个ID一行的结构,包含两个核心字段:

  • time_1对应的exposure值
  • time_3对应的outcome值(无time_3记录的ID填充NA)

示例数据(修正笔误后)

ID <- c(1,1,2,2,2,3,3,3,4,4)
exposure <-c(1.2, 1.3, 1.4, 1.5, 2.1, 2.2, 3.2, 4.2, 5.2, 6.2)
outcome <-c(0.8, 0.7, 0.6, 0.5, 0.4, 0.3, 0.2, 0.1, 2.1, 3.1)
Time<-c("time_1","time_2","time_1","time_2","time_3","time_1","time_2","time_3","time_1","time_2")
data <-data.frame(ID,exposure,outcome,Time)

方法一:分表提取+左连接(直观易理解)

通过分别提取time_1和time_3的记录,再用左连接保留所有ID,自动填充缺失值为NA:

library(tidyverse)

# 提取time_1的暴露数据
time1_exposure <- data %>%
  filter(Time == "time_1") %>%
  select(ID, exposure_time1 = exposure)

# 提取time_3的结局数据
time3_outcome <- data %>%
  filter(Time == "time_3") %>%
  select(ID, outcome_time3 = outcome)

# 左连接保留所有ID,无time_3记录的自动填充NA
final_data <- time1_exposure %>%
  left_join(time3_outcome, by = "ID")

运行后得到的结果:

> final_data
  ID exposure_time1 outcome_time3
1  1            1.2            NA
2  2            1.4           0.4
3  3            2.2           0.1
4  4            5.2            NA

方法二:pivot_wider直接重塑(更简洁)

利用pivot_wider直接将长格式转宽格式,通过参数控制缺失值填充:

library(tidyverse)

final_data <- data %>%
  # 仅保留需要的时间点
  filter(Time %in% c("time_1", "time_3")) %>%
  # 重塑宽格式,指定ID为分组,Time为列名,提取对应变量
  pivot_wider(
    id_cols = ID,
    names_from = Time,
    values_from = c(exposure, outcome),
    values_fill = NA  # 缺失值填充为NA
  ) %>%
  # 筛选并重命名需要的列
  select(ID, exposure_time1 = exposure_time_1, outcome_time3 = outcome_time_3)

此方法与方法一结果完全一致,适合熟悉tidyr语法的用户。


内容的提问来源于stack exchange,提问作者Aura

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.10 07:10:23