You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在R语言自然语言处理中保留指定停用词及查看停用词列表

R语言NLP预处理:保留"no"并查看移除的停用词

一、保留停用词中的"no"

要避免removeWords移除"no",只需先从英文停用词列表里剔除"no",再用修改后的列表执行移除操作:

# 获取英文停用词列表,剔除"no"
custom_stopwords <- stopwords('english')[stopwords('english') != 'no']

# 用自定义停用词列表执行移除操作
cleanset <- tm_map(corpus, removeWords, custom_stopwords)

完整预处理链修改后如下:

# Pre-processing chain
corpus <- tm_map(corpus, tolower)
corpus <- tm_map(corpus, removePunctuation)
corpus <- tm_map(corpus, removeNumbers)
# 自定义停用词列表,排除"no"
custom_stopwords <- stopwords('english')[stopwords('english') != 'no']
cleanset <- tm_map(corpus, removeWords, custom_stopwords) # 不再移除"no"
cleanset <- tm_map(cleanset, stemDocument)
cleanset <- tm_map(cleanset, stripWhitespace)
inspect(cleanset[1:25])

二、查看所有被移除的停用词

你可以通过两种方式查看被移除的停用词:

  • 直接输出原始英文停用词列表(即stopwords('english')),这就是默认会被移除的所有词,除了我们手动排除的"no":
# 查看默认英文停用词(所有会被移除的词,不含保留的"no")
print(stopwords('english'))
  • 如果想确认实际被移除的词,也可以对比处理前后的语料词汇,比如提取处理前的词汇集合,减去处理后的词汇集合:
# 提取处理前的词汇
pre_words <- unlist(strsplit(tolower(unlist(corpus)), "\\W+"))
pre_words <- unique(pre_words[pre_words != ""])

# 提取处理后的词汇
post_words <- unlist(strsplit(tolower(unlist(cleanset)), "\\W+"))
post_words <- unique(post_words[post_words != ""])

# 查看被移除的词
removed_words <- setdiff(pre_words, post_words)
print(removed_words)

内容的提问来源于stack exchange,提问作者wisamb

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.21 03:32:07