You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何检测DataFrame描述列中的单词是否属于指定词典?

嘿,我明白你的痛点了——要给每个商品描述做批量检查,只要有一个词在词典或自定义列表里就标记为TRUE,之前用qdapDictionaries和tokenizers没搞定多单词的情况对吧?我来给你捋个清晰的解决方案!

核心思路

我们需要把每个描述拆成独立单词,统一大小写后,逐个检查是否属于目标词典(内置词典+自定义词),最后对每一行判断“是否存在至少一个匹配词”。

完整代码示例

1. 加载所需库

先确保你安装了这些包,然后加载:

library(tidyverse)  # 用于数据处理(dplyr + stringr)
library(qdapDictionaries)  # 提供英文词典
library(tokenizers)  # 用于拆分单词

2. 准备你的数据

如果已经有现成的DataFrame可以跳过这步,这里先构造示例数据:

df <- tibble(
  Items = c("Item1", "Item2", "Item3", "Item4"),
  Descriptions = c("poster", "used cd music etc", "hckd herbal ingds.", "823942 blc")
)

3. 定义目标词典

结合qdap的内置词典和你的自定义单词,统一转成小写避免大小写匹配问题:

# 从qdap获取GradyAugmented常用英文单词库,转小写
base_dictionary <- tolower(GradyAugmented)

# 加入你的自定义单词(比如想把"etc"、"ingds."这类缩写也算进去)
custom_words <- c("etc", "ingds")
full_dictionary <- c(base_dictionary, custom_words)

注意:tokenize_words()会自动去掉标点,所以如果你的自定义词带标点,比如"ingds.",可以先把描述里的标点移除,或者直接把自定义词改成无标点的版本。

4. 批量检查并生成新列

这里用rowwise()让R逐行处理每个描述,拆分单词后做匹配:

df_result <- df %>%
  rowwise() %>%
  mutate(
    # 拆分描述为单词,自动去掉标点、转小写
    cleaned_tokens = list(tokenize_words(tolower(Descriptions))[[1]]),
    # 检查当前行是否有任意单词在词典里
    inDictionary = any(cleaned_tokens %in% full_dictionary)
  ) %>%
  ungroup() %>%
  select(-cleaned_tokens)  # 删掉中间生成的token列(可选)

5. 查看结果

运行后你会得到符合预期的输出:

print(df_result)
# # A tibble: 4 × 3
#   Items Descriptions       inDictionary
#   <chr> <chr>              <lgl>       
# 1 Item1 poster             TRUE        
# 2 Item2 used cd music etc  TRUE        
# 3 Item3 hckd herbal ingds. TRUE        
# 4 Item4 823942 blc         FALSE       

关键细节解释

  • 统一大小写:把描述和词典都转成小写,避免因为"Poster"和"poster"大小写不同导致匹配失败。
  • 自动处理标点:tokenize_words()会自动过滤掉描述里的标点符号,让单词匹配更准确。
  • 逐行处理:rowwise()确保每个描述的单词拆分和匹配都是独立进行的,不会把所有行的单词混在一起。

替代方案(不用rowwise)

如果你不习惯用rowwise(),可以用purrr::map_lgl()实现同样效果,代码更简洁,大数据量下速度也更快:

df_result <- df %>%
  mutate(
    inDictionary = map_lgl(Descriptions, ~ {
      # 对每个描述单独处理
      tokens <- tokenize_words(tolower(.x))[[1]]
      any(tokens %in% full_dictionary)
    })
  )

内容的提问来源于stack exchange,提问作者rued

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.08 14:07:34