如何在R的Quanteda Tokens中去除文本中的下划线
解决Quanteda处理Corpus转Tokens时的下划线残留与字符丢失问题
问题场景
使用R语言的Quanteda包将Corpus对象转换为Tokens时,遇到两类异常:
- 调用
tokens()时,内置参数无法去除单词/字符中的下划线,输出出现s_、_t这类带下划线的片段 - 用
stri_replace_all_regex()直接替换下划线后,原本带下划线的s、t等字符直接消失,不符合预期
原始代码及输出
dirty_corpus <- corpus(textdata) toks <- dirty_corpus %>% stringi::stri_replace_all_regex("'[a-z]*", "") %>% tokens(what = "word", remove_punct = TRUE, preserve_tags = FALSE, remove_numbers = TRUE, remove_separators = TRUE, remove_url = TRUE, split_hyphens = TRUE, remove_symbols = TRUE, split_tags = TRUE, verbose = TRUE) %>% tokens_remove(pattern = phrase(english_stopwords), valuetype = 'fixed') %>% tokens_wordstem() %>% tokens_tolower()
输出结果:
text6 : [1] "ys" "s_" "_t" "_s" "sw" "lnk" "smn" "pstd" "dwn" "blw" [11] "srri"
期望输出
text6 : [1] "ys" "s" "t" "s" "sw" "lnk" "smn" "pstd" "dwn" "blw" [11] "srri"
添加下划线替换后的代码及异常输出
dirty_corpus <- corpus(textdata) toks <- dirty_corpus %>% stringi::stri_replace_all_regex("'[a-z]*", "") %>% stringi::stri_replace_all_regex("_", "") %>% tokens(what = "word", remove_punct = TRUE, preserve_tags = FALSE, remove_numbers = TRUE, remove_separators = TRUE, remove_url = TRUE, split_hyphens = TRUE, remove_symbols = TRUE, split_tags = TRUE, verbose = TRUE) %>% tokens_remove(pattern = phrase(english_stopwords), valuetype = 'fixed') %>% tokens_wordstem() %>% tokens_tolower()
输出结果:
text6 : [1] "ys" "sw" "lnk" "smn" "pstd" "dwn" "blw" "srri"
问题根源
事后排查发现:替换下划线后生成的s、t这类单个字符,属于english_stopwords停用词列表中的条目,被后续tokens_remove()步骤直接过滤删除,并非下划线替换操作导致字符消失。
解决方案
针对该问题,有两种可行处理方式:
方案1:自定义停用词列表
从默认停用词中排除s、t这类需要保留的单个字符,再执行过滤:
# 生成排除指定字符的自定义停用词 custom_stopwords <- setdiff(english_stopwords, c("s", "t")) toks <- dirty_corpus %>% stringi::stri_replace_all_regex("'[a-z]*", "") %>% stringi::stri_replace_all_regex("_", "") %>% tokens(what = "word", remove_punct = TRUE, preserve_tags = FALSE, remove_numbers = TRUE, remove_separators = TRUE, remove_url = TRUE, split_hyphens = TRUE, remove_symbols = TRUE, split_tags = TRUE, verbose = TRUE) %>% tokens_remove(pattern = phrase(custom_stopwords), valuetype = 'fixed') %>% tokens_wordstem() %>% tokens_tolower()
方案2:调整处理顺序(推荐)
先完成Token化,再在Token层面替换下划线,避免提前生成单个停用词字符:
toks <- dirty_corpus %>% stringi::stri_replace_all_regex("'[a-z]*", "") %>% tokens(what = "word", remove_punct = TRUE, preserve_tags = FALSE, remove_numbers = TRUE, remove_separators = TRUE, remove_url = TRUE, split_hyphens = TRUE, remove_symbols = TRUE, split_tags = TRUE, verbose = TRUE) %>% # 在Token对象中直接替换下划线 tokens_replace(pattern = "_", replacement = "", valuetype = "fixed") %>% tokens_remove(pattern = phrase(english_stopwords), valuetype = 'fixed') %>% tokens_wordstem() %>% tokens_tolower()
总结
核心问题是替换下划线后生成的单个字符属于停用词范畴,被后续过滤步骤移除。通过调整停用词列表或处理顺序,即可得到保留s、t等字符的预期结果。
内容的提问来源于stack exchange,提问作者DartLazer
相关产品推荐
相关产品推荐

