如何在R语言数据框中统计每行的唯一匹配字符串模式数量
Count Unique Matching Patterns per Row in R
Got it, let's fix this for you! The original code counts the total number of times any of your target terms appear, but you want to count how many unique terms from toMatch show up at least once in each row. Here are two reliable approaches:
Approach 1: Using stringi with Row-wise Sum (Most Efficient)
This method leverages stri_detect_regex to check for the presence of each term, then sums the matches per row:
library(stringi) # Your original data setup dat <- read.table(text="index string 1 'I have first and second' 2 'I have first, first' 3 'I have second and first and thirdeen'", header=TRUE) toMatch <- c('first', 'second', 'third') # Create word-boundary regex patterns to avoid partial matches term_patterns <- paste0('\\b', toMatch, '\\b') # Detect each pattern across all rows, then sum matches per row dat$unique_match_count <- rowSums(sapply(term_patterns, function(pattern) { stri_detect_regex(dat$string, pattern) }))
How this works:
term_patternsadds word boundaries (\\b) to each term, ensuring we don't accidentally match partial words (like avoiding "thirdeen" being mistaken for "third").sapplyrunsstri_detect_regexfor each pattern, returning a logical matrix where each row corresponds to a row in your data, and each column corresponds to a term intoMatch(TRUE = term is present, FALSE = not present).rowSumsconverts those logical values to 1s and 0s, then sums them up to get the count of unique terms per row.
Approach 2: Using Word Extraction & Intersection (More Explicit)
If you prefer to explicitly extract words and check overlaps, this method works too (note: it requires handling punctuation carefully):
dat$unique_match_count <- sapply(dat$string, function(text) { # Extract all whole words (ignores punctuation attached to words) extracted_words <- stri_extract_all_regex(text, '\\b\\w+\\b')[[1]] # Count how many terms from toMatch are present in the extracted words length(intersect(extracted_words, toMatch)) })
Result Check
After running either method, your dat will look like this:
index string unique_match_count 1 1 'I have first and second' 2 2 2 'I have first, first' 1 3 3 'I have second and first and thirdeen' 2
Perfect! That's exactly the unique term count you wanted.
内容的提问来源于stack exchange,提问作者LMach
相关产品推荐
相关产品推荐

