技术问询:调整并行处理关键词统计函数,统计文档中出现的关键词个数
Got it, let's tweak your existing parallel function to track how many distinct keywords appear in each document instead of tallying total keyword occurrences. Here's a practical, efficient implementation in R that keeps the parallel processing capability intact:
Step 1: Define Your Input Data
First, let's confirm your sample data (I'll reuse your example):
keywords <- c("Toyota", "Prius", "BMW", "M3") documents <- c( "New Toyota Prius for sale, the Toyota Prius is in good condition", "BMW M3 that drives like a ..." )
Step 2: Parallel-Friendly Core Function
We'll create a function that checks which keywords are present in a single document, then counts how many unique ones show up. We'll use stringr for reliable string detection, and base R's parallel package for parallelization:
library(parallel) library(stringr) # Core function: Count unique keywords present in one document count_unique_keywords <- function(doc, keywords) { # Check each keyword's presence (adjust ignore_case to TRUE if needed) keyword_found <- str_detect(doc, regex(keywords, ignore_case = FALSE)) # Sum the number of keywords that were found sum(keyword_found) }
Step 3: Run Parallel Processing
Set up a cluster to leverage multiple CPU cores, then process all documents in parallel:
# Set up parallel cluster (leave one core free for system tasks) num_cores <- detectCores() - 1 cl <- makeCluster(num_cores) # Export required objects/functions to the cluster nodes clusterExport(cl, c("keywords", "count_unique_keywords", "str_detect", "regex")) # Process documents in parallel results <- parLapply(cl, documents, count_unique_keywords, keywords = keywords) # Clean up the cluster stopCluster(cl) # Format results for readability results_vector <- unlist(results) names(results_vector) <- paste0("Document_", seq_along(documents)) # View output print(results_vector)
Expected Output
Running this code will give you:
Document_1 Document_2 2 2
Which matches what we want: Document 1 has 2 unique keywords (Toyota, Prius), Document 2 has 2 (BMW, M3).
Optional Tweaks
- Case-Insensitive Matching: Change
ignore_case = FALSEtoTRUEin theregex()call if you want "toyota" or "PRIUS" to count as matches. - Exact Word Matching: To avoid partial matches (e.g., not counting "Priuses" as "Prius"), modify the regex to use word boundaries:
keyword_found <- str_detect(doc, regex(paste0("\\b", keywords, "\\b"), ignore_case = FALSE))
内容的提问来源于stack exchange,提问作者Tshabat

