如何在R中将同一列存储的两个变量值拆分到对应列?
在R中清洗拆分错位的数据
1. 先模拟你的数据
首先把表格数据转换成R能处理的数据框:
df <- data.frame( Column_A = c(1, 4, 7, 10), Column_B = c("2 3", "5", "8", "11 12"), Column_C = c(NA, "6", "9", ""), stringsAsFactors = FALSE )
2. 用tidyverse工具包处理(新手友好)
推荐用tidyverse,语法直观易读,先安装并加载包:
# 首次使用先安装 install.packages("tidyverse") library(tidyverse)
然后执行清洗步骤:
df_clean <- df %>% # 拆分Column B,按空格分成两列,不足的补NA separate(Column_B, into = c("B1", "B2"), sep = " ", fill = "right") %>% # 把B2的值填充到Column C的空/NA位置,原C有值的保留 mutate(Column_C = ifelse(is.na(Column_C) | Column_C == "", B2, Column_C)) %>% # 重命名B1为原Column B,删除临时列B2 rename(Column_B = B1) %>% select(-B2) %>% # 把所有列转成数值型 mutate(across(everything(), as.numeric))
运行后得到的干净数据:
> df_clean Column_A Column_B Column_C 1 1 2 3 2 4 5 6 3 7 8 9 4 10 11 12
3. 基础R实现(无需额外包)
如果不想装包,用基础R代码也能完成:
# 拆分Column B的字符串 split_b <- strsplit(df$Column_B, " ") # 更新Column B为拆分后的第一个数值 df$Column_B <- sapply(split_b, function(x) x[1]) # 提取拆分后的第二个数值,填充到Column C的空值位置 b2_vals <- sapply(split_b, function(x) if(length(x) >= 2) x[2] else NA) df$Column_C <- ifelse(is.na(df$Column_C) | df$Column_C == "", b2_vals, df$Column_C) # 转换所有列为数值型 df[] <- lapply(df, as.numeric)
内容的提问来源于stack exchange,提问作者Hassan Nasir Mirbahar
相关产品推荐
相关产品推荐

