You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用R语言tidytext包正确移除停用词?解决特殊停用词未移除问题

问题原因与解决方法

核心原因:字符编码不匹配

你遇到的问题是因为输入文本中的撇号是智能引号(弯撇:’),而stop_words数据集中对应的停用词用的是直撇('),这两个是完全不同的Unicode字符,所以%in%匹配时会判定为不同字符串,导致过滤失败。

可以用以下代码验证差异:

# 查看stop_words中的目标停用词
stop_words %>% filter(word %in% c("don't", "it's", "i've"))

# 对比两种撇号的编码差异
charToRaw("don’t")  # 输出 b'e2 80 99',对应弯撇
charToRaw("don't")  # 输出 b'27',对应直撇

解决办法:统一撇号格式

只需要把输入文本中的弯撇替换成直撇,就能和stop_words中的停用词匹配上,具体代码如下:

library(tidyverse)
library(tidytext)

data(stop_words)
example_words <- c("the", "quick", "brown", "fox", "jumps", "over", "the", "lazy", "dog","i'm","don’t","it’s","i’ve")

# 将所有弯撇替换为直撇,统一字符格式
example_words_clean <- str_replace_all(example_words, "’", "'")

# 执行过滤
filtered_words <- example_words_clean[!example_words_clean %in% stop_words$word]
filtered_words

执行后输出结果:

[1] "quick" "brown" "fox"   "jumps" "lazy"  "dog"

这类问题通常是复制富文本(比如Word文档、网页文本)时,编辑器自动将直撇转换为智能引号导致的,统一字符格式就能解决。

内容的提问来源于stack exchange,提问作者student_R123

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.25 14:25:08