在R语言自然语言处理中保留指定停用词及查看停用词列表
R语言NLP预处理:保留"no"并查看移除的停用词
一、保留停用词中的"no"
要避免removeWords移除"no",只需先从英文停用词列表里剔除"no",再用修改后的列表执行移除操作:
# 获取英文停用词列表,剔除"no" custom_stopwords <- stopwords('english')[stopwords('english') != 'no'] # 用自定义停用词列表执行移除操作 cleanset <- tm_map(corpus, removeWords, custom_stopwords)
完整预处理链修改后如下:
# Pre-processing chain corpus <- tm_map(corpus, tolower) corpus <- tm_map(corpus, removePunctuation) corpus <- tm_map(corpus, removeNumbers) # 自定义停用词列表,排除"no" custom_stopwords <- stopwords('english')[stopwords('english') != 'no'] cleanset <- tm_map(corpus, removeWords, custom_stopwords) # 不再移除"no" cleanset <- tm_map(cleanset, stemDocument) cleanset <- tm_map(cleanset, stripWhitespace) inspect(cleanset[1:25])
二、查看所有被移除的停用词
你可以通过两种方式查看被移除的停用词:
- 直接输出原始英文停用词列表(即
stopwords('english')),这就是默认会被移除的所有词,除了我们手动排除的"no":
# 查看默认英文停用词(所有会被移除的词,不含保留的"no") print(stopwords('english'))
- 如果想确认实际被移除的词,也可以对比处理前后的语料词汇,比如提取处理前的词汇集合,减去处理后的词汇集合:
# 提取处理前的词汇 pre_words <- unlist(strsplit(tolower(unlist(corpus)), "\\W+")) pre_words <- unique(pre_words[pre_words != ""]) # 提取处理后的词汇 post_words <- unlist(strsplit(tolower(unlist(cleanset)), "\\W+")) post_words <- unique(post_words[post_words != ""]) # 查看被移除的词 removed_words <- setdiff(pre_words, post_words) print(removed_words)
内容的提问来源于stack exchange,提问作者wisamb
相关产品推荐
相关产品推荐

