如何将数据框中father_details列的邮箱与父亲姓名对应到同一行?
长格式数据集合并父亲姓名与邮箱的实现思路
核心思路
你的数据集按每4行对应一个用户的长格式存储:第一行是用户基本信息+父亲姓名,第四行是父亲邮箱,中间两行是无效NA行。处理步骤如下:
- 分组:按每4行划分为一个用户组,确保同一用户的所有信息在同一组内。
- 提取邮箱:在每组内筛选出
father_details中包含@的内容(即父亲邮箱)。 - 合并行:将提取到的父亲邮箱赋值给该组内有用户姓名的行(每组第一行),生成
father_email列。 - 过滤无效行:只保留每组内有用户姓名的有效行,删除其余NA行。
R代码实现(dplyr版本)
library(dplyr) # 加载样本数据 df <- structure(list(Name = c("John", NA, NA, NA, "Chris", NA, NA, NA), Email = c("john@abc.com", NA, NA, NA, "chris@abc.com", NA, NA, NA), father_details = c("herald", NA, NA, "herald@abc.com", "matthew", NA, NA, "matthew@abc.com")), class = "data.frame", row.names = c(NA, 8L)) # 数据处理 result <- df %>% # 每4行分为一组 mutate(group = gl(n()/4, 4)) %>% group_by(group) %>% # 提取组内的父亲邮箱(含@的非NA值) mutate(father_email = first(na.omit(father_details[grepl("@", father_details)]))) %>% # 仅保留有用户姓名的有效行 filter(!is.na(Name)) %>% # 移除分组辅助列 select(-group) print(result)
代码解释
gl(n()/4, 4):根据总行数自动生成分组标识,确保每4行归为一组。first(na.omit(father_details[grepl("@", father_details)])):在组内筛选出邮箱格式的内容,去除NA后取唯一的邮箱值。filter(!is.na(Name)):过滤掉所有无用户姓名的NA行,只保留有效数据行。
R代码实现(base R版本)
如果不想依赖dplyr包,也可以用基础R实现:
# 加载样本数据 df <- structure(list(Name = c("John", NA, NA, NA, "Chris", NA, NA, NA), Email = c("john@abc.com", NA, NA, NA, "chris@abc.com", NA, NA, NA), father_details = c("herald", NA, NA, "herald@abc.com", "matthew", NA, NA, "matthew@abc.com")), class = "data.frame", row.names = c(NA, 8L)) # 生成分组标识 df$group <- rep(1:(nrow(df)/4), each = 4) # 按分组处理数据 result <- do.call(rbind, lapply(split(df, df$group), function(grp) { # 提取当前组的父亲邮箱 father_email <- grp$father_details[grepl("@", grp$father_details) & !is.na(grp$father_details)] # 保留组内有用户姓名的行,并添加father_email列 valid_row <- grp[!is.na(grp$Name), ] valid_row$father_email <- father_email return(valid_row) })) # 移除分组列 result$group <- NULL print(result)
内容的提问来源于stack exchange,提问作者Ramakrishna S
相关产品推荐
相关产品推荐

