You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

技术问询:调整并行处理关键词统计函数,统计文档中出现的关键词个数

Count Unique Keywords Per Document (Parallel Processing)

Got it, let's tweak your existing parallel function to track how many distinct keywords appear in each document instead of tallying total keyword occurrences. Here's a practical, efficient implementation in R that keeps the parallel processing capability intact:

Step 1: Define Your Input Data

First, let's confirm your sample data (I'll reuse your example):

keywords <- c("Toyota", "Prius", "BMW", "M3")
documents <- c(
  "New Toyota Prius for sale, the Toyota Prius is in good condition",
  "BMW M3 that drives like a ..."
)

Step 2: Parallel-Friendly Core Function

We'll create a function that checks which keywords are present in a single document, then counts how many unique ones show up. We'll use stringr for reliable string detection, and base R's parallel package for parallelization:

library(parallel)
library(stringr)

# Core function: Count unique keywords present in one document
count_unique_keywords <- function(doc, keywords) {
  # Check each keyword's presence (adjust ignore_case to TRUE if needed)
  keyword_found <- str_detect(doc, regex(keywords, ignore_case = FALSE))
  # Sum the number of keywords that were found
  sum(keyword_found)
}

Step 3: Run Parallel Processing

Set up a cluster to leverage multiple CPU cores, then process all documents in parallel:

# Set up parallel cluster (leave one core free for system tasks)
num_cores <- detectCores() - 1
cl <- makeCluster(num_cores)

# Export required objects/functions to the cluster nodes
clusterExport(cl, c("keywords", "count_unique_keywords", "str_detect", "regex"))

# Process documents in parallel
results <- parLapply(cl, documents, count_unique_keywords, keywords = keywords)

# Clean up the cluster
stopCluster(cl)

# Format results for readability
results_vector <- unlist(results)
names(results_vector) <- paste0("Document_", seq_along(documents))

# View output
print(results_vector)

Expected Output

Running this code will give you:

Document_1 Document_2 
         2          2 

Which matches what we want: Document 1 has 2 unique keywords (Toyota, Prius), Document 2 has 2 (BMW, M3).

Optional Tweaks

  • Case-Insensitive Matching: Change ignore_case = FALSE to TRUE in the regex() call if you want "toyota" or "PRIUS" to count as matches.
  • Exact Word Matching: To avoid partial matches (e.g., not counting "Priuses" as "Prius"), modify the regex to use word boundaries:
    keyword_found <- str_detect(doc, regex(paste0("\\b", keywords, "\\b"), ignore_case = FALSE))
    

内容的提问来源于stack exchange,提问作者Tshabat

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 10:15:48