You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用quanteda字典创建关键词列并丢弃含嵌套目标词的长匹配

问题1:引用DFM列时需要小写键名的原因

quanteda在构建dfm(文档特征矩阵)时默认会将所有特征名自动转换为小写形式,即便你定义字典时键名写的是大写IE,最终生成的dfm列名也会被标准化为小写ie,因此引用时必须使用小写形式才能匹配到对应列。
如果希望保留原大写键名,可以在调用dfm()时添加参数tolower = FALSE:

dfm <- tokens_lookup(toks, dictionary = dict, nested_scope = "dictionary", case_insensitive = F) %>%
  tokens_remove("Northern Ireland") %>% 
  dfm(tolower = FALSE)
# 此时就可以用大写键名引用
df$contains <- as.logical(dfm[, "IE"], case_insensitive = FALSE)

问题2:kwic排除指定字典键的实现方法

有两种常用方案可以实现仅保留Ireland的匹配结果:

方案1:调用kwic时直接指定需要匹配的字典键

不需要传入完整字典,仅提取你需要的IE键对应的匹配规则即可:

# 仅匹配IE键对应的词汇
words <- kwic(toks, pattern = dict["IE"], case_insensitive = FALSE)

方案2:生成kwic结果后过滤不需要的键

如果已经生成了全量匹配结果,可以直接过滤掉不需要的关键词行:

words <- kwic(toks, pattern = dict, case_insensitive = FALSE) %>%
  dplyr::filter(keyword != "Northern Ireland")

处理多匹配结果时你现有的pivot_wider逻辑是正确的,先过滤不需要的匹配项再做宽表转换,就能得到仅含Ireland匹配结果的关键词列。

内容的提问来源于stack exchange,提问作者Jasper

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.03 21:57:04