使用Tidyverse判断数据框字符串列是否含指定向量元素并标记
Here's how you can achieve this using tidyverse tools (dplyr for data manipulation and stringr for string operations):
Step 1: Load Required Packages
First, load the tidyverse package which includes both dplyr and stringr:
library(tidyverse)
Step 2: Define Your Data
Your original data frame and check vector:
df <- data.frame(ID = 1:5, string = c("blah, F21, blah", "woop, woop, F25", "G1, yes, yes", "hey, hey F23", "how, G2, how")) check <- c("F21", "F23", "G1")
Step 3: Create the Matching Pattern
Collapse the check vector into a regex pattern where elements are separated by | (acts as an OR operator). Adding word boundaries (\\b) ensures we match whole elements only, avoiding unintended partial matches (e.g., preventing "F210" from being mistaken for "F21"):
match_pattern <- str_c("\\b", check, "\\b", collapse = "|")
Step 4: Add the Test Column
Use mutate() and if_else() to create the test column based on whether the string contains any element from the check vector:
df <- df %>% mutate(test = if_else(str_detect(string, match_pattern), "in check", "not in check"))
Result
The resulting data frame will match your expected output:
print(df)
Output:
ID string test 1 1 blah, F21, blah in check 2 2 woop, woop, F25 not in check 3 3 G1, yes, yes in check 4 4 hey, hey F23 in check 5 5 how, G2, how not in check
Explanation
str_c()combines the check elements into a single regex pattern that looks for any of the target values.str_detect()checks each string to see if it matches any part of the pattern.if_else()assigns the appropriate label ("in check" or "not in check") based on the detection result.
内容的提问来源于stack exchange,提问作者Scott Purvis

