R中使用dplyr拆分tibble字符串并保留匹配目标条目的最优方法
最优实现方案
无需生成中间列、无需行遍历/拆分行操作,单步mutate即可得到目标结果,天然兼容多等位基因突变的多匹配场景:
library(tidyverse) d %>% mutate( final = str_extract_all(to_sep, "want (\\d+)", group = 1) %>% map_chr(~ if (length(.x) == 0) NA_character_ else paste(.x, collapse = ", ")) )
输出结果和预期一致,所有原有列完整保留:
# A tibble: 4 × 4 id to_sep other_info final <dbl> <chr> <chr> <chr> 1 1 want 543;discard;toss important 543 2 2 nothing;worth;keeping stuff NA 3 3 want 129 is 129 4 4 meh;want 657 here too 657
如果是多匹配的输入(如"want 123; want 456; other"),final列会自动拼接为"123, 456",可根据需要调整paste的分隔符。
符合purrr+case_when要求的实现
如果希望明确用case_when处理NA逻辑,可以用以下写法:
d %>% mutate( final = str_split(to_sep, ";") %>% map_chr(~ { match_nums <- str_subset(.x, "want") %>% str_extract("\\d+") case_when( length(match_nums) == 0 ~ NA_character_, TRUE ~ paste(match_nums, collapse = ", ") ) }) )
该写法逻辑和原有思路一致,但省略了行遍历、中间列生成、拆分组装的冗余步骤。
内容的提问来源于stack exchange,提问作者GenesRus
相关产品推荐
相关产品推荐

