首次文本挖掘中无法使用stop_words处理文本的技术咨询
文本挖掘代码问题修正与步骤指导
原代码存在的核心问题
- 文本拆分错误:
text1是单个长字符串,无法直接生成16行的数据框,需先按换行分割为每行元素 - 未保留分词结果:
unnest_tokens执行后未赋值给变量,后续无法基于分词数据处理 - 停用词过滤逻辑错误:
anti_join的目标应为stop_words,而非原数据框,且需指定匹配列 - 词频统计语法错误:
count函数需传入列名,而非整个数据框
修正后的完整代码
library(tm) library(tidytext) library(dplyr) library(ggplot2) # 原始文本,按换行分割为16行的向量 text1 <- str_split(c("Dear land of Guyana, of rivers and plains, Made rich by the sunshine, and lush by the rains, Set gem-like and fair between mounts and sea- Your children salute you. dear land of the free. Green land of Guyana, our heroes of yore, Both bondsman and free, laid their bones on your shore, This soil so they hallowed, and from them are we, All sons of one mother, Guyana the free Great land of Guyana, diverse though our strains, We are born of their sacrifice, heirs of their pains, And ours is the glory their eyes did not see – One Land of six peoples, united and free. Dear Land of Guyana, to you will we give Our homage, our service each day that we live; God guard you, great Mother, and make us to be More worthy our heritage – land of the free."), "\n")[[1]] # 创建包含行号和对应文本的数据框 newtext1 <- data.frame(line = 1:16, text = text1) # 分词处理,保存结果到分词数据框 tidy_text <- newtext1 %>% unnest_tokens(word, text) # 加载停用词数据集 data(stop_words) # 过滤停用词 tidy_text_no_stop <- tidy_text %>% anti_join(stop_words, by = "word") # 统计词频并排序 word_counts <- tidy_text_no_stop %>% count(word, sort = TRUE) # 查看结果 word_counts
关键步骤说明
- 用
str_split将原始长文本按换行符分割为16个元素的向量,确保数据框行号与文本行一一对应 - 分词后将结果存入
tidy_text,后续所有处理基于此数据框 anti_join通过by = "word"指定匹配列,过滤掉停用词count(word, sort = TRUE)统计每个词的出现次数并按降序排列
内容的提问来源于stack exchange,提问作者Rohan Sagar
相关产品推荐
相关产品推荐

