R语言:编写函数匹配列表中各DataFrame的值
给你的DataFrame列表写个值匹配函数的方案
先把你给出的DataFrame列表代码贴出来确认下:
a <- data.frame( Data0=c("Y","Y","Y","Y","Y","Y","N","N","N","N","N","N"), Data1=c(16,18,19,20,21,50,16,18,19,20,21,50), Data2=c(2.2291,2.0743,1.9369,1.8148,1.7064,1.6102,2.2291,2.0743,1.9369,1.8148,1.7064,1.6102) ) b <- data.frame( Data0=c(-2 , 0 , 1 , 2 , 3 , 4 , 5 , 6 , 7 , 8 , 9 ,10 ,11), Data1=c(0.8891 ,0.8891,0.9051,1,0.8891,0.8891,0.7907,0.8891,0.9929,0.8891,0.8891,0.8891,0.8891) ) dfl <- list(a,b)
接下来我给你写几个实用的函数,覆盖常见的值匹配场景,你可以根据自己的需求调整:
场景1:找特定值在每个DataFrame里的位置
这个函数会遍历列表里的每个DataFrame,返回目标值所在的行和列位置,方便你定位:
match_value_in_list <- function(df_list, target_value) { lapply(df_list, function(df) { # 用which找到匹配位置,arr.ind=TRUE返回行列索引 match_pos <- which(df == target_value, arr.ind = TRUE) # 转成DataFrame看起来更清晰 as.data.frame(match_pos) }) }
用起来很简单
比如你要找0.8891这个值的位置:
result <- match_value_in_list(dfl, 0.8891) # 查看第一个DataFrame的匹配结果 result[[1]] # 查看第二个DataFrame的匹配结果 result[[2]]
场景2:根据某一列的值匹配另一列的内容
比如你想根据Data1列的值,对应拿到Data0列的内容,这个函数就适合:
match_column_pair <- function(df_list, match_col, target_col) { lapply(df_list, function(df) { # 先检查指定的列是否存在,避免报错 if (!all(c(match_col, target_col) %in% colnames(df))) { stop("你指定的列在这个DataFrame里找不到哦!") } # 返回以match_col为名称、target_col为值的向量,方便匹配调用 setNames(df[[target_col]], df[[match_col]]) }) }
示例用法
比如在第一个DataFrame里,用Data1的值匹配Data0:
col_result <- match_column_pair(dfl, "Data1", "Data0") # 第一个DataFrame的匹配结果 col_result[[1]] # 第二个DataFrame的匹配结果 col_result[[2]]
要是需要模糊匹配怎么办?
如果你的需求是模糊匹配(比如找包含某个字符的内容),可以用这个版本:
fuzzy_match_in_list <- function(df_list, pattern, ignore_case = TRUE) { lapply(df_list, function(df) { # 只对字符型列做模糊匹配 char_cols <- sapply(df, is.character) lapply(df[, char_cols], function(col) { which(grepl(pattern, col, ignore.case = ignore_case)) }) }) }
这样你可以根据自己的实际需求选对应的函数,或者在这些基础上修改逻辑就行啦!
内容的提问来源于stack exchange,提问作者Helen
相关产品推荐
相关产品推荐

