You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在R的Quanteda Tokens中去除文本中的下划线

解决Quanteda处理Corpus转Tokens时的下划线残留与字符丢失问题

问题场景

使用R语言的Quanteda包将Corpus对象转换为Tokens时,遇到两类异常:

  1. 调用tokens()时,内置参数无法去除单词/字符中的下划线,输出出现s_、_t这类带下划线的片段
  2. 用stri_replace_all_regex()直接替换下划线后,原本带下划线的s、t等字符直接消失,不符合预期

原始代码及输出

dirty_corpus <- corpus(textdata)

toks <- dirty_corpus %>%
  stringi::stri_replace_all_regex("'[a-z]*", "") %>%
  tokens(what = "word", remove_punct = TRUE, preserve_tags = FALSE, remove_numbers = TRUE, remove_separators = TRUE,
         remove_url = TRUE, split_hyphens = TRUE, remove_symbols = TRUE, split_tags = TRUE, verbose = TRUE) %>%
  tokens_remove(pattern = phrase(english_stopwords), valuetype = 'fixed') %>%
  tokens_wordstem() %>%
  tokens_tolower()

输出结果:

text6 : [1] "ys"   "s_"   "_t"   "_s"   "sw"   "lnk"  "smn"  "pstd" "dwn"  "blw"  [11] "srri"

期望输出

text6 : [1] "ys"   "s"   "t"   "s"   "sw"   "lnk"  "smn"  "pstd" "dwn"  "blw"  [11] "srri"

添加下划线替换后的代码及异常输出

dirty_corpus <- corpus(textdata)

toks <- dirty_corpus %>%
  stringi::stri_replace_all_regex("'[a-z]*", "") %>%
  stringi::stri_replace_all_regex("_", "") %>%
  tokens(what = "word", remove_punct = TRUE, preserve_tags = FALSE, remove_numbers = TRUE, remove_separators = TRUE,
         remove_url = TRUE, split_hyphens = TRUE, remove_symbols = TRUE, split_tags = TRUE, verbose = TRUE) %>%
  tokens_remove(pattern = phrase(english_stopwords), valuetype = 'fixed') %>%
  tokens_wordstem() %>%
  tokens_tolower()

输出结果:

text6 : [1] "ys"   "sw"   "lnk"  "smn"  "pstd" "dwn"  "blw"  "srri"

问题根源

事后排查发现:替换下划线后生成的s、t这类单个字符,属于english_stopwords停用词列表中的条目,被后续tokens_remove()步骤直接过滤删除,并非下划线替换操作导致字符消失。

解决方案

针对该问题,有两种可行处理方式:

方案1:自定义停用词列表

从默认停用词中排除s、t这类需要保留的单个字符,再执行过滤:

# 生成排除指定字符的自定义停用词
custom_stopwords <- setdiff(english_stopwords, c("s", "t"))

toks <- dirty_corpus %>%
  stringi::stri_replace_all_regex("'[a-z]*", "") %>%
  stringi::stri_replace_all_regex("_", "") %>%
  tokens(what = "word", remove_punct = TRUE, preserve_tags = FALSE, remove_numbers = TRUE, remove_separators = TRUE,
         remove_url = TRUE, split_hyphens = TRUE, remove_symbols = TRUE, split_tags = TRUE, verbose = TRUE) %>%
  tokens_remove(pattern = phrase(custom_stopwords), valuetype = 'fixed') %>%
  tokens_wordstem() %>%
  tokens_tolower()

方案2:调整处理顺序(推荐)

先完成Token化,再在Token层面替换下划线,避免提前生成单个停用词字符:

toks <- dirty_corpus %>%
  stringi::stri_replace_all_regex("'[a-z]*", "") %>%
  tokens(what = "word", remove_punct = TRUE, preserve_tags = FALSE, remove_numbers = TRUE, remove_separators = TRUE,
         remove_url = TRUE, split_hyphens = TRUE, remove_symbols = TRUE, split_tags = TRUE, verbose = TRUE) %>%
  # 在Token对象中直接替换下划线
  tokens_replace(pattern = "_", replacement = "", valuetype = "fixed") %>%
  tokens_remove(pattern = phrase(english_stopwords), valuetype = 'fixed') %>%
  tokens_wordstem() %>%
  tokens_tolower()

总结

核心问题是替换下划线后生成的单个字符属于停用词范畴,被后续过滤步骤移除。通过调整停用词列表或处理顺序,即可得到保留s、t等字符的预期结果。

内容的提问来源于stack exchange,提问作者DartLazer

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.08 09:35:22