R语言数据重塑问题:将夫妻配对宽数据转换为指定结构
问题:重塑对偶数据,拆分自身与伴侣得分
原始数据集
初始的宽格式数据如下:
data <- data.frame( T2_husband = rnorm(5), T2_wife = rnorm(5), T1_husband = rnorm(5), T1_wife = rnorm(5), Dyad_ID = 1:5 )
目标结构
需要将数据转换为长格式,包含以下字段:
- Dyad_ID:对偶ID
- Role:角色(丈夫/妻子)
- My_T1:自身的T1得分
- My_T2:自身的T2得分
- Partner_T1:伴侣的T1得分(丈夫对应妻子的T1,妻子对应丈夫的T1)
- Partner_T2:伴侣的T2得分
尝试的代码(未达预期)
# Reshape the dataset reshaped_data <- data %>% pivot_longer( cols = c(T1_husband, T1_wife, T2_husband, T2_wife), # Columns to reshape names_to = c("Time", "Role"), # Split column names into "Time" and "Role" names_sep = "_" # Separator is "_" ) %>% pivot_wider( id_cols = c(Dyad_ID, Role), # Keep Dyad_ID and Role as unique identifiers names_from = Time, # Reshape Time into separate columns values_from = value # Values go into "T1" and "T2" columns ) # Add My_* and Partner_* columns reshaped_data <- reshaped_data %>% group_by(Dyad_ID) %>% # Group by Dyad_ID to match partner roles mutate( My_T1 = T1, My_T2 = T2, Partner_T1 = T1[Role != first(Role)], # Select T1 where Role is not the current Role Partner_T2 = T2[Role != first(Role)] # Select T2 where Role is not the current Role ) %>% ungroup() %>% # Remove grouping select(Dyad_ID, Role, My_T1, My_T2, Partner_T1, Partner_T2)
问题分析
代码前半部分的pivot_longer+pivot_wider逻辑正确,已经得到了每个角色的T1/T2得分。但后半部分获取伴侣得分的逻辑存在漏洞:first(Role)依赖组内角色的排序,如果某对偶组内角色顺序不是先丈夫后妻子,就会导致取值错误。比如组内先出现wife时,丈夫的Partner_T1会误取自己的T1值,完全不符合需求。
修正后的代码
我们可以针对当前角色明确指定伴侣角色的得分,逻辑更稳定:
library(tidyverse) # 第一步:重塑为每个角色一行的格式 reshaped_data <- data %>% pivot_longer( cols = -Dyad_ID, names_to = c("Time", "Role"), names_sep = "_" ) %>% pivot_wider( id_cols = c(Dyad_ID, Role), names_from = Time, values_from = value ) # 第二步:添加自身与伴侣得分 final_data <- reshaped_data %>% group_by(Dyad_ID) %>% mutate( My_T1 = T1, My_T2 = T2, # 根据当前角色匹配伴侣的得分 Partner_T1 = case_when( Role == "husband" ~ T1[Role == "wife"], Role == "wife" ~ T1[Role == "husband"] ), Partner_T2 = case_when( Role == "husband" ~ T2[Role == "wife"], Role == "wife" ~ T2[Role == "husband"] ) ) %>% ungroup() %>% select(Dyad_ID, Role, My_T1, My_T2, Partner_T1, Partner_T2)
也可以用更简洁的lead/lag写法(前提是每个对偶组恰好有两行):
final_data <- reshaped_data %>% group_by(Dyad_ID) %>% mutate( My_T1 = T1, My_T2 = T2, Partner_T1 = lag(T1) %>% replace_na(lead(T1)), Partner_T2 = lag(T2) %>% replace_na(lead(T2)) ) %>% ungroup() %>% select(Dyad_ID, Role, My_T1, My_T2, Partner_T1, Partner_T2)
效果验证
运行代码后,每个对偶组会生成两行数据:一行是丈夫(包含自身T1/T2和妻子的T1/T2),一行是妻子(包含自身T1/T2和丈夫的T1/T2),完全符合目标结构。
内容的提问来源于stack exchange,提问作者Tube
相关产品推荐
相关产品推荐

