You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在R语言数据框中统计每行的唯一匹配字符串模式数量

Count Unique Matching Patterns per Row in R

Got it, let's fix this for you! The original code counts the total number of times any of your target terms appear, but you want to count how many unique terms from toMatch show up at least once in each row. Here are two reliable approaches:

Approach 1: Using stringi with Row-wise Sum (Most Efficient)

This method leverages stri_detect_regex to check for the presence of each term, then sums the matches per row:

library(stringi)

# Your original data setup
dat <- read.table(text="index string
1 'I have first and second'
2 'I have first, first'
3 'I have second and first and thirdeen'", header=TRUE)
toMatch <- c('first', 'second', 'third')

# Create word-boundary regex patterns to avoid partial matches
term_patterns <- paste0('\\b', toMatch, '\\b')

# Detect each pattern across all rows, then sum matches per row
dat$unique_match_count <- rowSums(sapply(term_patterns, function(pattern) {
  stri_detect_regex(dat$string, pattern)
}))

How this works:

  • term_patterns adds word boundaries (\\b) to each term, ensuring we don't accidentally match partial words (like avoiding "thirdeen" being mistaken for "third").
  • sapply runs stri_detect_regex for each pattern, returning a logical matrix where each row corresponds to a row in your data, and each column corresponds to a term in toMatch (TRUE = term is present, FALSE = not present).
  • rowSums converts those logical values to 1s and 0s, then sums them up to get the count of unique terms per row.

Approach 2: Using Word Extraction & Intersection (More Explicit)

If you prefer to explicitly extract words and check overlaps, this method works too (note: it requires handling punctuation carefully):

dat$unique_match_count <- sapply(dat$string, function(text) {
  # Extract all whole words (ignores punctuation attached to words)
  extracted_words <- stri_extract_all_regex(text, '\\b\\w+\\b')[[1]]
  # Count how many terms from toMatch are present in the extracted words
  length(intersect(extracted_words, toMatch))
})

Result Check

After running either method, your dat will look like this:

index                          string unique_match_count
1     1        'I have first and second'                  2
2     2              'I have first, first'                  1
3     3 'I have second and first and thirdeen'                  2

Perfect! That's exactly the unique term count you wanted.

内容的提问来源于stack exchange,提问作者LMach

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 08:22:52