在tibble与tidytext中使用stop_words时遇报错求助
问题解决:
ncol(corpus) >= 2 is not TRUE 报错处理 错误原因
你的word_space已经是经过unnest_tokens拆分后的单列单词数据框,但你用tibble(text = word_space)把它打包成了嵌套数据列(一列里包含另一个数据框),unnest_tokens无法处理这种输入结构,因此抛出该错误。
解决方案
方案一:直接基于现有word_space处理
word_space已经是拆分好的单词列表,无需再次执行unnest_tokens,直接执行去停用词+统计即可:
word_space %>% anti_join(stop_words) %>% count(word, sort = TRUE)
方案二:从原始数据cam重新完整处理
如果希望从原始数据开始走完整流程,建议合并步骤,避免中间多余的count操作导致后续混乱:
cam %>% unnest_tokens(input = value, output = word, token = "words") %>% anti_join(stop_words) %>% count(word, sort = TRUE)
内容的提问来源于stack exchange,提问作者user18454095
相关产品推荐
相关产品推荐

