使用R生成词云时丢失特定词汇AI的问题求助
问题根源
词云丢失"AI"的核心原因:
- 预处理时
tolower()将"AI"转为小写"ai" stopwords("english")默认包含古英语代词"ai"(意为"我"),导致小写后的"ai"被当作停用词移除
解决方案
修改停用词列表,从默认英文停用词中剔除"ai",再结合自定义停用词使用:
# 读取数据等原有步骤保持不变... # 调整停用词处理逻辑 # 1. 获取默认英文停用词并移除"ai" default_stopwords <- setdiff(stopwords("english"), "ai") # 2. 定义自定义停用词 my_stopwords <- c("learning", "development", "inperson", "streamed", "live", "design", "training", "new", "better", "can", "beyond", "like", "know", "without", "almost", "don't") # 3. 合并为最终停用词列表 final_stopwords <- c(default_stopwords, my_stopwords) # 预处理文本时使用新的停用词列表 corpus <- tm_map(corpus, content_transformer(tolower)) corpus <- tm_map(corpus, removePunctuation) corpus <- tm_map(corpus, removeNumbers) corpus <- tm_map(corpus, removeWords, final_stopwords) # 后续词云生成步骤保持不变...
验证方法
单独测试第83行文本,确认"ai"被保留:
test_corpus <- Corpus(VectorSource(titles[83])) test_corpus <- tm_map(test_corpus, content_transformer(tolower)) test_corpus <- tm_map(test_corpus, removePunctuation) test_corpus <- tm_map(test_corpus, removeNumbers) test_corpus <- tm_map(test_corpus, removeWords, final_stopwords) # 查看处理后的文本,应包含"ai" test_corpus[[1]][1]
若希望词云显示大写的"AI",可在预处理后添加替换步骤:
corpus <- tm_map(corpus, content_transformer(function(x) gsub("ai", "AI", x)))
内容的提问来源于stack exchange,提问作者Simon Shin
相关产品推荐
相关产品推荐

