You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何高效同时统计含不同特定字符串的数据行数量?

高效统计数据框单列中多组特定字符串的匹配行数

当需要统计数据框单列中分别包含多个特定字符串的行数量时,无需逐个重复编写sum(grepl(...))或sum(str_count(...)),以下是几种高效的批量处理方法:

方法1:基础R的sapply批量处理

通过定义目标字符串向量,结合sapply遍历执行匹配与求和,代码简洁且高效:

# 定义需要统计的目标字符串集合
target_strings <- c("great", "utr", "DISTAL", "LONG")

# 批量计算每个字符串对应的匹配行数
count_result <- sapply(target_strings, function(pattern) {
  sum(grepl(pattern, q.data$string))
})

# 转成直观的数据框格式(可选)
count_df <- data.frame(
  target_string = names(count_result),
  matched_rows = count_result,
  row.names = NULL
)

方法2:Tidyverse风格(stringr + purrr)

如果习惯使用tidyverse工具链,可通过map系列函数实现批量统计:

library(stringr)
library(purrr)

target_strings <- c("great", "utr", "DISTAL", "LONG")

count_df <- map_dbl(target_strings, ~sum(str_detect(q.data$string, .x))) %>%
  set_names(target_strings) %>%
  enframe(name = "target_string", value = "matched_rows")

方法3:向量化匹配矩阵(大数据量最优)

针对超大数据集,一次性生成匹配矩阵再按列求和,能减少重复匹配的开销:

target_strings <- c("great", "utr", "DISTAL", "LONG")

# 生成匹配逻辑矩阵:行=数据行,列=目标字符串
match_matrix <- sapply(target_strings, grepl, x = q.data$string)

# 按列求和得到每个字符串的匹配行数
count_result <- colSums(match_matrix)

以上方法只需维护target_strings向量即可扩展任意数量的统计目标,避免了重复代码,且执行效率远高于逐个编写求和语句的基础方法。

内容的提问来源于stack exchange,提问作者Charité Learner

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.16 07:53:08