求助:使用R语言quanteda包统计客户评论停用词数量的保存问题
Hey there! I get it—you’ve got your text tokenized and filtered out the stop words, but now you need to tally how many stop words exist in each individual comment. Let’s fix that quickly with a couple of straightforward approaches using quanteda.
Method 1: Count Stop Words Directly from Your Tokens
If you haven’t already saved the filtered stop words, you can do this in one go by keeping only stop words and then counting their lengths per document:
# Load quanteda if you haven't already library(quanteda) # Assume corpus2.tokens is your pre-tokenized text object # Get the default stop words list (swap in a custom list if needed) stop_words_list <- stopwords() # Keep only stop words and count their number per document stopword_counts <- tokens_select( corpus2.tokens, pattern = stop_words_list, selection = "keep", valuetype = "fixed" ) %>% lengths() # Convert to a readable data frame result_df <- data.frame( comment_id = names(stopword_counts), stopword_count = as.numeric(stopword_counts), stringsAsFactors = FALSE ) # Preview the result head(result_df)
Method 2: Use Your Existing Filtered Stop Words Object
If you already have the stop.words object from your code, you can just count the length of tokens in each document directly:
# Count stop words per document from your existing stop.words tokens stopword_counts <- lengths(stop.words) # Convert to a structured data frame result_df <- data.frame( comment_id = names(stopword_counts), stopword_count = as.numeric(stopword_counts), stringsAsFactors = FALSE )
Bonus: Merge with Comment Metadata
If your original corpus has document metadata (like the full comment text, user IDs, timestamps, etc.), you can merge it into your result to add context:
# Assuming corpus2 is your original corpus with docvars library(dplyr) result_df <- result_df %>% left_join(docvars(corpus2), by = c("comment_id" = "docname"))
Save the Final Result
Once you have your structured table, save it to a CSV or other format for further analysis:
write.csv(result_df, "stopword_counts_per_comment.csv", row.names = FALSE)
This will give you a clean, usable dataset where each row represents a single comment with its corresponding stop word count—exactly what you need!
内容的提问来源于stack exchange,提问作者Ceri

