R语言情感分析:停用词未移除及清理结果无法保存问题求助
解决R语言情感分析里的停用词&输出保存坑
兄弟,我太懂你这种做情感分析时停用词删不掉、处理好的结果存不住的痛苦了!这些都是R文本处理里的常见坑,我帮你逐个拆解解决:
一、先搞定内置英文停用词不生效的问题
十有八九是你预处理的顺序错了!举个用tm包的正确流程,你对照着改:
library(tm) # 加载你的正负向推文语料库 corpus <- VCorpus(DirSource("你的语料文件夹路径")) # 预处理步骤一定要按这个顺序来! corpus <- tm_map(corpus, content_transformer(tolower)) # 先转全小写,不然停用词匹配不上 corpus <- tm_map(corpus, removePunctuation) # 移除标点 corpus <- tm_map(corpus, removeNumbers) # 移除数字 # 关键一步:移除内置停用词 corpus <- tm_map(corpus, removeWords, stopwords("english"))
很多人踩坑是先分词再转小写,导致停用词(都是小写)和文本里的大写单词匹配不上,顺序绝对不能乱!
二、用外部stopwords.txt的正确姿势
从GitHub下的停用词txt,得先转成R能识别的字符向量再用:
# 读取停用词文件,注意路径要对 custom_stopwords <- readLines("stopwords.txt") # 清理掉空行(txt文件里可能有) custom_stopwords <- custom_stopwords[custom_stopwords != ""] # 用这个自定义停用词清理语料 corpus <- tm_map(corpus, removeWords, custom_stopwords)
要是你用tidytext(现在很多人用这个更灵活),流程会更清爽:
library(tidytext) library(dplyr) # 把语料转成 tidy 格式 tidy_tweets <- corpus %>% tidy() %>% unnest_tokens(word, text) %>% # 先删内置停用词,再删自定义的 anti_join(stop_words, by = "word") %>% anti_join(tibble(word = custom_stopwords), by = "word")
三、处理后输出存不住?这么办!
不管是tm的语料库还是tidytext的 tidy 数据,保存方法都很直接:
1. 保存tm处理后的语料库
# 存成R专属的rds文件,下次直接加载不用再预处理 saveRDS(corpus, file = "processed_tweets_corpus.rds") # 或者导出成纯文本文件 writeLines(sapply(corpus, as.character), "processed_tweets.txt")
别直接用write.table存VCorpus对象,会乱码,一定要用sapply把每个文档转成字符再导出!
2. 保存tidy格式的处理结果
# 存成csv,方便用Excel看 write.csv(tidy_tweets, "processed_tweets_tidy.csv", row.names = FALSE) # 或者存成rds,保留R的数据结构 saveRDS(tidy_tweets, "processed_tweets_tidy.rds")
四、词云前的最后检查
生成对比词云前,先查一下高频词,确认停用词真的被删掉了:
# 用tm包生成词频矩阵 dtm <- DocumentTermMatrix(corpus) word_freq <- colSums(as.matrix(dtm)) # 看前20个高频词 head(sort(word_freq, decreasing = TRUE), 20)
要是这里还能看到"the""and"这类停用词,回去检查预处理顺序或者停用词向量是不是没加载对。
内容的提问来源于stack exchange,提问作者Mahnoor
相关产品推荐
相关产品推荐

