You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用R语言检索百万行CSV文本关键词并生成累计得分列时的data.table语法错误问询

Fixing the := Error in R when Matching Keywords and Calculating Cumulative Scores

Error Cause

The error Error in := (id, 1L) : Check that is.data.table(DT) == TRUE happens because you're trying to use data.table's exclusive := assignment operator on a character vector (your corpus object after na.omit()), not a data.table.

Additionally, there are a couple of other issues in your preprocessing steps:

  • tm_map(corpus, removeNumbers) requires corpus to be a Corpus object from the tm package, but you converted it to a character vector earlier, which would trigger another error.
  • You lost the connection between your cleaned text and the original dataset when you extracted corpus as a standalone vector, making it hard to map scores back to the original rows.

Corrected Solution with Full Code

Here's a complete, fixed workflow that handles text preprocessing correctly, matches keywords, and calculates cumulative scores while keeping all data aligned:

1. Load Required Packages

First, make sure you have these packages installed and loaded:

library(data.table)
library(tm)
library(stringr)

2. Import and Preprocess Data

We'll work directly with the original CNN dataset to preserve row relationships:

# Import the CSV (disable string factors to keep text as strings)
CNN <- read.csv("CNNNewsDataSet.csv", stringsAsFactors = FALSE)

# Inspect data structure
str(CNN)

# Clean the message column step-by-step
CNN$clean_message <- CNN$message %>%
  # Keep only printable ASCII characters
  gsub('[^\\x20-\\x7E]', '', .) %>%
  # Replace empty strings with NA
  ifelse(. == "", NA, .) %>%
  # Remove punctuation and extra whitespace
  gsub("[[:punct:][:blank:]]+", " ", .) %>%
  # Convert all text to lowercase
  tolower() %>%
  # Remove numbers
  gsub("[0-9]+", "", .)

# Optional: Remove rows with empty cleaned messages
CNN <- CNN[!is.na(CNN$clean_message), ]

3. Define Keywords and Their Scores

We'll lowercase keywords to match our cleaned text:

search_for <- data.table(
  word = tolower(c("Capitol", "Biden", "Congress", "Marines", "Senate", "White House")),
  value = c(-0.5, -0.6, -0.4, -0.2, -0.4, -0.04)
)

4. Match Keywords and Calculate Cumulative Scores

Option 1: Using row-wise detection (simpler for large datasets)

# Convert CNN to data.table for efficient operations
setDT(CNN)

# Loop through each keyword to mark matches and calculate individual scores
for (i in seq_len(nrow(search_for))) {
  current_word <- search_for$word[i]
  current_score <- search_for$value[i]
  
  # Create a column to flag if the word is present
  CNN[, paste0("match_", current_word) := str_detect(clean_message, fixed(current_word))]
  # Calculate score contribution for this word
  CNN[, paste0("score_", current_word) := ifelse(get(paste0("match_", current_word)), current_score, 0)]
}

# Compute total cumulative score
score_columns <- grep("score_", colnames(CNN), value = TRUE)
CNN[, total_score := rowSums(.SD), .SDcols = score_columns]

# Collect all matched keywords for each row
match_columns <- grep("match_", colnames(CNN), value = TRUE)
CNN[, matched_words := apply(.SD, 1, function(row) {
  paste(search_for$word[row], collapse = ", ")
}), .SDcols = match_columns]

# View the final result (original message, cleaned text, matched words, total score)
final_result <- CNN[, .(message, clean_message, matched_words, total_score)]
print(final_result)

Option 2: Using Cartesian join (closer to your original approach)

# Create a data.table with cleaned text and row IDs
text_dt <- data.table(
  row_id = seq_len(nrow(CNN)),
  text = CNN$clean_message
)

# Perform Cartesian join with keyword table
search_res <- merge(text_dt[, id := 1L], search_for[, id := 1L], by = "id", allow.cartesian = TRUE)

# Check if each keyword matches the text
search_res[, match := str_detect(text, fixed(word))]

# Aggregate scores and matched words per row
search_res_agg <- search_res[match == TRUE, .(
  matched_words = paste(sort(word), collapse = ", "),
  total_score = sum(value)
), by = .(row_id, text)]

# Merge back to original dataset
final_result <- merge(CNN, search_res_agg, by.x = "row_id", by.y = "row_id", all.x = TRUE)

# Fill NA values for rows with no matches
final_result[, `:=`(
  matched_words = ifelse(is.na(matched_words), "", matched_words),
  total_score = ifelse(is.na(total_score), 0, total_score)
)]

print(final_result)

内容的提问来源于stack exchange,提问作者Michael71

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.30 22:12:30