使用R的tidyverse按3列规整透视表格,匹配访视与对应年龄
用tidyverse实现根据visit匹配对应age列
问题背景
现有数据集包含user_ID、age_v1/age_v2/age_v3、gender、visit和score列,每行对应一次访问记录,需要根据visit列的取值(v1/v2/v3)提取对应的年龄值,生成新的age列,同时避免使用pivot_longer导致数据膨胀。
解决方案
以下提供两种tidyverse实现方式,按需选择:
方法1:固定匹配(适合已知有限的visit类型)
通过case_when直接匹配visit值与对应age列:
library(tidyverse) # 构造示例数据 df <- tibble( user_ID = c(1, 1, 1), age_v1 = c(63, 63, 63), age_v2 = c(65, 65, 65), age_v3 = c(67, 67, 67), gender = c("M", "M", "M"), visit = c("v1", "v2", "v3"), score = c(8, 4, 1) ) # 处理数据 df_processed <- df %>% mutate(age = case_when( visit == "v1" ~ age_v1, visit == "v2" ~ age_v2, visit == "v3" ~ age_v3, TRUE ~ NA_real_ # 处理未匹配的异常情况 )) %>% select(user_ID, age, gender, visit, score) # 保留目标列
方法2:动态匹配(适合扩展更多visit类型)
通过字符串替换生成对应age列名,动态提取值,无需修改匹配逻辑:
df_processed <- df %>% mutate(age = cur_data()[[str_replace(visit, "v", "age_v")]]) %>% select(user_ID, age, gender, visit, score)
输出结果
# A tibble: 3 × 5 user_ID age gender visit score <dbl> <dbl> <chr> <chr> <dbl> 1 1 63 M v1 8 2 1 65 M v2 4 3 1 67 M v3 1
内容的提问来源于stack exchange,提问作者pdw5
相关产品推荐
相关产品推荐

