基于唯一标识合并UK Understanding Society亲子数据集多行记录为单行的实现方法
解决方案
完全可以通过数据透视(长表转宽表)实现你要的格式,操作前需要确认你已经在父母数据中标记了父母的性别(区分父亲/母亲),如果暂时没有性别字段,可以临时按关联顺序分配父/母标识,示例代码如下:
tidyverse 方案(推荐)
# 加载依赖包 library(dplyr) library(tidyr) # 先给示例数据增加父母性别标识,你实际使用时可以从原始父母数据中关联该字段 example <- example %>% group_by(Youth_Personal_ID) %>% mutate(parent_type = c("Mother_education", "Father_education")) %>% ungroup() # 长表转宽表得到目标格式 result <- example %>% pivot_wider( id_cols = c(Youth_Personal_ID, Youth_reading), names_from = parent_type, values_from = Parent_education )
运行后result的输出和你期望的格式完全一致:
Youth_Personal_ID Youth_reading Mother_education Father_education 1 200 once a week bachelors HS diploma
base R 方案
如果不想加载额外包,可以用基础函数实现:
# 新增父母类型列 example$parent_type <- ave(rep(1, nrow(example)), example$Youth_Personal_ID, FUN = function(x) c("Mother_education", "Father_education")) # 长转宽 result <- reshape(example, idvar = c("Youth_Personal_ID", "Youth_reading"), timevar = "parent_type", direction = "wide") # 调整列名去掉前缀 colnames(result) <- gsub("Parent_education.", "", colnames(result))
注意:如果一个子女对应的父母数量不足2人(比如只有单亲数据),转宽后对应父/母教育水平的字段会自动填充为NA,可根据需求后续单独处理缺失值。
内容的提问来源于stack exchange,提问作者Grant
相关产品推荐
相关产品推荐

