使用R语言检索百万行CSV文本关键词并生成累计得分列时的data.table语法错误问询
:= Error in R when Matching Keywords and Calculating Cumulative Scores Error Cause
The error Error in := (id, 1L) : Check that is.data.table(DT) == TRUE happens because you're trying to use data.table's exclusive := assignment operator on a character vector (your corpus object after na.omit()), not a data.table.
Additionally, there are a couple of other issues in your preprocessing steps:
tm_map(corpus, removeNumbers)requirescorpusto be aCorpusobject from thetmpackage, but you converted it to a character vector earlier, which would trigger another error.- You lost the connection between your cleaned text and the original dataset when you extracted
corpusas a standalone vector, making it hard to map scores back to the original rows.
Corrected Solution with Full Code
Here's a complete, fixed workflow that handles text preprocessing correctly, matches keywords, and calculates cumulative scores while keeping all data aligned:
1. Load Required Packages
First, make sure you have these packages installed and loaded:
library(data.table) library(tm) library(stringr)
2. Import and Preprocess Data
We'll work directly with the original CNN dataset to preserve row relationships:
# Import the CSV (disable string factors to keep text as strings) CNN <- read.csv("CNNNewsDataSet.csv", stringsAsFactors = FALSE) # Inspect data structure str(CNN) # Clean the message column step-by-step CNN$clean_message <- CNN$message %>% # Keep only printable ASCII characters gsub('[^\\x20-\\x7E]', '', .) %>% # Replace empty strings with NA ifelse(. == "", NA, .) %>% # Remove punctuation and extra whitespace gsub("[[:punct:][:blank:]]+", " ", .) %>% # Convert all text to lowercase tolower() %>% # Remove numbers gsub("[0-9]+", "", .) # Optional: Remove rows with empty cleaned messages CNN <- CNN[!is.na(CNN$clean_message), ]
3. Define Keywords and Their Scores
We'll lowercase keywords to match our cleaned text:
search_for <- data.table( word = tolower(c("Capitol", "Biden", "Congress", "Marines", "Senate", "White House")), value = c(-0.5, -0.6, -0.4, -0.2, -0.4, -0.04) )
4. Match Keywords and Calculate Cumulative Scores
Option 1: Using row-wise detection (simpler for large datasets)
# Convert CNN to data.table for efficient operations setDT(CNN) # Loop through each keyword to mark matches and calculate individual scores for (i in seq_len(nrow(search_for))) { current_word <- search_for$word[i] current_score <- search_for$value[i] # Create a column to flag if the word is present CNN[, paste0("match_", current_word) := str_detect(clean_message, fixed(current_word))] # Calculate score contribution for this word CNN[, paste0("score_", current_word) := ifelse(get(paste0("match_", current_word)), current_score, 0)] } # Compute total cumulative score score_columns <- grep("score_", colnames(CNN), value = TRUE) CNN[, total_score := rowSums(.SD), .SDcols = score_columns] # Collect all matched keywords for each row match_columns <- grep("match_", colnames(CNN), value = TRUE) CNN[, matched_words := apply(.SD, 1, function(row) { paste(search_for$word[row], collapse = ", ") }), .SDcols = match_columns] # View the final result (original message, cleaned text, matched words, total score) final_result <- CNN[, .(message, clean_message, matched_words, total_score)] print(final_result)
Option 2: Using Cartesian join (closer to your original approach)
# Create a data.table with cleaned text and row IDs text_dt <- data.table( row_id = seq_len(nrow(CNN)), text = CNN$clean_message ) # Perform Cartesian join with keyword table search_res <- merge(text_dt[, id := 1L], search_for[, id := 1L], by = "id", allow.cartesian = TRUE) # Check if each keyword matches the text search_res[, match := str_detect(text, fixed(word))] # Aggregate scores and matched words per row search_res_agg <- search_res[match == TRUE, .( matched_words = paste(sort(word), collapse = ", "), total_score = sum(value) ), by = .(row_id, text)] # Merge back to original dataset final_result <- merge(CNN, search_res_agg, by.x = "row_id", by.y = "row_id", all.x = TRUE) # Fill NA values for rows with no matches final_result[, `:=`( matched_words = ifelse(is.na(matched_words), "", matched_words), total_score = ifelse(is.na(total_score), 0, total_score) )] print(final_result)
内容的提问来源于stack exchange,提问作者Michael71

