使用tidymodels的textrecipes处理tokenlist时拼写检查报错如何解决
解决方案
错误根因
step_tokenize() 处理后的输出是 textrecipes 专属的 tokenlist 嵌套列表结构,而非普通字符向量,无法直接传入要求字符输入的 hunspell 系列函数,因此触发 is.character(words) is not TRUE 报错。
修正方案
使用 textrecipes 内置的 step_tokenapply() 函数,它专门用于对 tokenlist 中的每个分词向量执行自定义变换,处理后仍保留 tokenlist 结构,不会影响后续NLP预处理流程。
调整后完整可运行代码
library(tidyverse) library(tidymodels) library(textrecipes) library(hunspell) product_descriptions <- tibble( desc = c("goood product", "not sou good", "vad produkt"), price = c(1000, 700, 250) ) correct_spelling <- function(input) { output <- case_when( # 检查拼写,必要时校正 !hunspell_check(input, dictionary('en_US')) ~ hunspell_suggest(input, dictionary('en_US')) %>% # 取第一个推荐结果,列表为空则返回NA map(1, .default = NA) %>% unlist(), TRUE ~ input # 单词拼写正确则直接返回 ) # 拼写错误但无推荐结果则返回原单词 ifelse(is.na(output), input, output) } # 仅需将原来的step_mutate替换为step_tokenapply即可 product_recipe <- recipe(desc ~ price, data = product_descriptions) %>% step_tokenize(desc) %>% step_tokenapply(desc, correct_spelling) # 运行预处理 prepped_rec <- product_recipe %>% prep() # 查看校正后的分词结果 prepped_rec %>% juice()
效果验证
如果需要查看校正后的完整文本,可以在流程末尾增加step_untokenize步骤:
recipe(desc ~ price, data = product_descriptions) %>% step_tokenize(desc) %>% step_tokenapply(desc, correct_spelling) %>% step_untokenize(desc) %>% prep() %>% juice()
运行后将得到拼写校正后的文本列:
- "good product"
- "not so good"
- "bad product"
完全匹配你不使用recipes时的处理效果。
内容的提问来源于stack exchange,提问作者dzegpi
相关产品推荐
相关产品推荐

