R语言基于首末数字匹配跨数据框修改列值及删除行问题
R语言字符串匹配替换与行过滤实现方案
首先预设变量:第一个数据框为df1,第二个数据框为df2,两个数据框中存储4段数字分隔字符串的列名为str_col,可根据实际情况替换为你自己的列名。
tidyverse 实现方案(语法直观,适合新手)
- 加载依赖包
library(tidyverse)
- 构造df1的匹配对照表
# 拆分出首尾数字作为匹配键,去重避免多匹配 df1_map <- df1 %>% separate( col = str_col, into = c("first", "n2", "n3", "last"), sep = "-", remove = FALSE ) %>% distinct(first, last, .keep_all = TRUE) %>% select(str_col, first, last)
- 处理df2完成需求
df2_output <- df2 %>% # 拆分df2的字符串取首尾匹配键 separate( col = str_col, into = c("df2_first", "tmp1", "tmp2", "df2_last"), sep = "-", remove = FALSE ) %>% # 关联匹配表 left_join(df1_map, by = c("df2_first" = "first", "df2_last" = "last")) %>% # 过滤无匹配的行 filter(!is.na(str_col.y)) %>% # 替换为df1的对应值 mutate(str_col = str_col.y) %>% # 保留df2原始的列结构 select(all_of(colnames(df2)))
原生base R实现方案(无需额外安装包)
# 生成df1的匹配向量,键为"首数字_尾数字",值为对应的原字符串 df1_keys <- paste( sapply(strsplit(df1$str_col, "-"), function(x) x[1]), sapply(strsplit(df1$str_col, "-"), function(x) x[4]), sep = "_" ) match_vec <- setNames(df1$str_col, df1_keys) # 处理df2 df2_keys <- paste( sapply(strsplit(df2$str_col, "-"), function(x) x[1]), sapply(strsplit(df2$str_col, "-"), function(x) x[4]), sep = "_" ) df2$str_col <- match_vec[df2_keys] # 剔除无匹配的行 df2_output <- df2[!is.na(df2$str_col), ]
适配调整说明
- 如果字符串不是固定4段分隔,仅需要取首尾段匹配,将上述代码中取拆分后第4位的部分修改为
x[length(x)]即可适配任意长度的分隔字符串 - 若df1中存在多组首尾数字相同的字符串,上述代码默认取首次出现的字符串作为替换值,如有其他优先级规则,可提前对df1按规则排序后再生成匹配表即可
内容的提问来源于stack exchange,提问作者Anna
相关产品推荐
相关产品推荐

