如何在R中使用星号通配符匹配并统计数据框内目标词的出现频次
问题原因
你当前使用==做精确全等匹配,不会识别*的通配符语义,因此带星号的条目无法匹配到符合前缀规则的单词,统计结果不符合预期。另外还要注意区分带通配符和不带通配符的匹配逻辑,避免不带通配符的条目出现模糊匹配的问题。
修复代码
Items <- c('decid*','head', 'heads') df1<-data.frame(Items) words<- c('head', 'heads', 'decided', 'decides', 'top', 'undecided') df_main<-data.frame(words) item <- vector() count <- vector() for (i in 1:length(unique(Items))){ item[i] <- Items[i] if (grepl("\\*", item[i])) { # 带通配符的条目:将*转为正则的任意匹配规则,锚定前缀避免匹配到中间含对应字符的单词 pattern <- sub("\\*", ".*", paste0("^", item[i])) count[i] <- sum(grepl(pattern, df_main$words)) } else { # 不带通配符的条目:保留原精确匹配逻辑 count[i] <- sum(df_main$words == item[i]) } } word_freq <- data.frame(cbind(item, count)) word_freq
输出结果
运行上述代码后即可得到你期望的统计结果:
| item | count |
|---|---|
| decid* | 2 |
| head | 1 |
| heads | 1 |
补充说明
如果你的通配符不需要限制前缀匹配(即允许匹配单词任意位置包含对应字符的内容,比如decid*也统计undecided),只需要去掉paste0("^", item[i])中的^锚定符即可。
内容的提问来源于stack exchange,提问作者Asghar
相关产品推荐
相关产品推荐

