使用anti_join去除停用词报错,求R语言替代解决方法
解决anti_join去除停用词的报错问题
错误原因
你的报错核心是stop_words未被识别为标准数据框,反而变成了一串正则字符串,导致与你的数据框x数据源不匹配。大概率是未正确加载tidytext包内置的stop_words数据集,或是误操作将其转换成了字符串类型。
解决方案
方案1:确保正确加载tidytext并指定join列
先确认已加载tidytext包,然后在anti_join中明确指定匹配列,必要时加上copy=TRUE参数兼容数据源差异:
library(tidytext) library(dplyr) library(tidyr) data_review2 <- data %>% unnest_tokens(word, reviewText) %>% anti_join(stop_words, by = "word", copy = TRUE)
方案2:用filter+%in%替代anti_join
如果join方式仍有问题,直接用filter结合%in%过滤停用词,逻辑更直观且不易出数据源匹配问题:
library(tidytext) library(dplyr) library(tidyr) data_review2 <- data %>% unnest_tokens(word, reviewText) %>% filter(!word %in% stop_words$word)
方案3:手动定义停用词列表(备用)
如果内置stop_words仍无法正常调用,可手动创建停用词向量直接使用:
# 示例英文停用词列表,可根据需求扩展 custom_stop_words <- c("i", "me", "my", "myself", "we", "our", "ours", "ourselves", "you", "your", "yours", "yourself", "yourselves") data_review2 <- data %>% unnest_tokens(word, reviewText) %>% filter(!word %in% custom_stop_words)
内容的提问来源于stack exchange,提问作者Mayank Kumar Singh
相关产品推荐
相关产品推荐

