使用TryCatch无法捕获right_join警告指定的报错行问题求助
问题:捕获right_join警告对应的问题行失败
执行
right_join时出现如下警告:Warning message: In right_join(., strat_dat, by = "idnr") : Each row in `x` is expected to match at most 1 row in `y`. ℹ Row 4987 of `x` matches multiple rows. ℹ If multiple matches are expected, set `multiple = "all"` to silence this warning.
以下是示例代码:
library(dplyr) # 创建示例数据框 df_bm <- data.frame(a = c(1, 2, 3), b = c("1", "2", "3")) strat_dat <- data.frame(idnr = c(11,22,22,33), mem = c("ja", "nee", "ja", "nee")) # 捕获错误并存储问题行的函数 catch_error <- function() { tryCatch( { df_bm <- df_bm %>% mutate(idnr = as.numeric(paste0(b, a)), .before = a) %>% right_join(strat_dat, by = "idnr") %>% filter(mem == "ja") %>% select(-c(a, b, mem)) }, warning = function(w) { message(paste("Warning:", w$message)) rows <- which(duplicated(df_bm$idnr)) rows_out <<- rows message(paste("Problematic rows:", toString(rows))) } ) } # 运行函数 catch_error()
运行后输出:
Warning: Each row in `x` is expected to match at most 1 row in `y`. Problematic rows:
请问我哪里操作错误?
问题分析与修正
核心错误点
操作对象错误
你在警告处理函数里用的是原始的df_bm,但触发警告的是经过mutate添加了idnr列后的数据集。原始df_bm根本没有idnr列,所以duplicated(df_bm$idnr)会返回全FALSE,自然找不到任何问题行。误解警告含义
这个警告不是说x(处理后的df_bm)里有重复的idnr,而是x中的某一行在y(strat_dat)中匹配到了多个条目。实际是strat_dat里存在重复的idnr(比如示例里的22出现两次),当x的行匹配到这个idnr时,就会触发警告。你需要找的是x中哪些idnr在y里有多个匹配,而非x自身的重复行。
修正后的代码
library(dplyr) df_bm <- data.frame(a = c(1, 2, 3), b = c("1", "2", "3")) strat_dat <- data.frame(idnr = c(11,22,22,33), mem = c("ja", "nee", "ja", "nee")) catch_error <- function() { tryCatch( { # 先处理数据,保存中间结果,避免覆盖原始数据 df_processed <- df_bm %>% mutate(idnr = as.numeric(paste0(b, a)), .before = a) # 找出y中重复的idnr duplicate_ids <- strat_dat$idnr[duplicated(strat_dat$idnr)] %>% unique() # 定位x中匹配到这些重复idnr的行 problematic_rows <- which(df_processed$idnr %in% duplicate_ids) # 执行后续的连接和筛选操作 df_result <- df_processed %>% right_join(strat_dat, by = "idnr") %>% filter(mem == "ja") %>% select(-c(a, b, mem)) # 返回结果和问题行,方便后续使用 list(result = df_result, problematic_rows = problematic_rows) }, warning = function(w) { message(paste("Warning:", w$message)) # 在警告处理中重新生成处理后的数据集,定位问题行 df_processed <- df_bm %>% mutate(idnr = as.numeric(paste0(b, a)), .before = a) duplicate_ids <- strat_dat$idnr[duplicated(strat_dat$idnr)] %>% unique() problematic_rows <- which(df_processed$idnr %in% duplicate_ids) rows_out <<- problematic_rows message(paste("Problematic rows:", toString(problematic_rows))) } ) } # 运行函数 catch_error()
修正说明
- 先保存
mutate后的中间数据df_processed,确保操作的是触发警告的正确数据集。 - 先找出
strat_dat里重复的idnr,再匹配df_processed中对应的行,这些就是触发警告的问题行。 - 避免直接覆盖原始的
df_bm,改用中间变量让逻辑更清晰,也方便后续调试。
内容的提问来源于stack exchange,提问作者Tom
相关产品推荐
相关产品推荐

