You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Tidymodels textrecipes+spacyR分词后如何去除标点?

解决spacyR分词词形还原后去除标点/数字token的方法

你可以使用textrecipes包中的step_filter_tokens()步骤,在词形还原后过滤掉标点、数字这类不需要的token。具体实现如下:

修改后的代码

library(tidyverse)
library(tidymodels)
library(textrecipes)
library(spacyr)

text = "It was a day, Tuesday. It wasn't Thursday!"

df <- tibble(text)

spacyr::spacy_initialize(entity = FALSE)

lexicon_features_tokenized_lemmatised <-
  recipe(~ text, data = df%>%head(1)) %>%
  step_tokenize(text, engine = "spacyr") %>%
  step_lemma(text) %>%
  # 添加过滤步骤:移除标点符号
  step_filter_tokens(text, filter = function(token) {
    # 匹配非标点的token并保留
    !stringr::str_detect(token, "^[[:punct:]]$")
  }) %>%
  prep() %>%
  bake(new_data = NULL) 

lexicon_features_tokenized_lemmatised %>% pull(text) %>% textrecipes:::get_tokens()

代码说明

  • step_filter_tokens()的filter参数接受自定义函数,对每个token做判断:返回TRUE保留token,返回FALSE移除token。
  • 这里用stringr::str_detect(token, "^[[:punct:]]$")匹配单个标点符号(正则[:punct:]覆盖所有标点类字符),取反后即可移除所有单个标点token。
  • 如果还需要移除数字,可修改过滤条件:
    filter = function(token) {
      !stringr::str_detect(token, "^[[:punct:]]$") & !stringr::str_detect(token, "^[[:digit:]]$")
    }
    

运行修改后的代码,输出即为你期望的结果:"it", "be", "a", "day", "Tuesday", "it", "be", "not", "Thursday"

内容的提问来源于stack exchange,提问作者GeorgeM

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.25 23:12:15