You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用tidymodels的textrecipes处理tokenlist时拼写检查报错如何解决

解决方案

错误根因

step_tokenize() 处理后的输出是 textrecipes 专属的 tokenlist 嵌套列表结构,而非普通字符向量,无法直接传入要求字符输入的 hunspell 系列函数,因此触发 is.character(words) is not TRUE 报错。

修正方案

使用 textrecipes 内置的 step_tokenapply() 函数,它专门用于对 tokenlist 中的每个分词向量执行自定义变换,处理后仍保留 tokenlist 结构,不会影响后续NLP预处理流程。

调整后完整可运行代码

library(tidyverse)
library(tidymodels)
library(textrecipes)
library(hunspell)

product_descriptions <- tibble(
  desc = c("goood product", "not sou good", "vad produkt"),
  price = c(1000, 700, 250)
)

correct_spelling <- function(input) {
  output <- case_when(
    # 检查拼写,必要时校正
    !hunspell_check(input, dictionary('en_US')) ~
      hunspell_suggest(input, dictionary('en_US')) %>%
      # 取第一个推荐结果,列表为空则返回NA
      map(1, .default = NA) %>%
      unlist(),
    TRUE ~ input # 单词拼写正确则直接返回
  )
  # 拼写错误但无推荐结果则返回原单词
  ifelse(is.na(output), input, output)
}

# 仅需将原来的step_mutate替换为step_tokenapply即可
product_recipe <- recipe(desc ~ price, data = product_descriptions) %>% 
  step_tokenize(desc) %>% 
  step_tokenapply(desc, correct_spelling)

# 运行预处理
prepped_rec <- product_recipe %>% prep()

# 查看校正后的分词结果
prepped_rec %>% juice()

效果验证

如果需要查看校正后的完整文本,可以在流程末尾增加step_untokenize步骤:

recipe(desc ~ price, data = product_descriptions) %>% 
  step_tokenize(desc) %>% 
  step_tokenapply(desc, correct_spelling) %>%
  step_untokenize(desc) %>% 
  prep() %>% 
  juice()

运行后将得到拼写校正后的文本列:

  • "good product"
  • "not so good"
  • "bad product"
    完全匹配你不使用recipes时的处理效果。

内容的提问来源于stack exchange,提问作者dzegpi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.26 12:36:04