You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用spacyr词形还原时如何保留词间连字符?

问题描述

我用spacyr对演讲语料库做词形还原,之后用quanteda分词,再借助textstat_frequency()分析结果。当前遇到的核心问题:

  • 直接用quanteda分词时,带词间连字符的关键术语会被保留为单个词元,完全符合预期;
  • 但先经过spacyr词形还原后,带连字符的词汇无法保持整体结构。

试过nounphrase_consolidate()方法,虽能保留部分连字符词汇,但一致性极差——有时目标术语会被单独保留,有时却被合并进更大的名词短语,没法满足后续用特定词典匹配分析的需求。已知SpaCy有相关解决方案,想咨询spacyr是否有类似选项。

附上相关代码(分词时设置remove_punct参数对结果无影响):

test.sp <- spacy_parse(test.corpus, lemma = TRUE, entity = FALSE, pos = FALSE, tag = FALSE, nounphrase = TRUE)
test.sp$token <- test.sp$lemma
test.np <- nounphrase_consolidate(test.sp)
test.tokens.3 <- as.tokens(test.np)
test.tokens.3 <- tokens(test.tokens.3, remove_symbols = TRUE,
                      remove_numbers = TRUE,
                      remove_punct = TRUE,
                      remove_url = TRUE) %>% 
  tokens_tolower() %>% 
  tokens_select(pattern = stopwords("en"), selection = "remove") 
解决方案

1. 给SpaCy词汇表添加自定义复合词规则

spacyr可以通过修改底层SpaCy的词汇属性,强制保留指定连字符术语的完整性:

  • 提前将目标连字符术语添加到SpaCy词汇表,标记为不可分复合词,这样词形还原时就不会拆分它们。
# 初始化spacy并加载模型
spacy_initialize(model = "en_core_web_sm")
# 获取SpaCy词汇表对象
vocab <- spacy_get_nlp()$vocab
# 定义需要保留的连字符术语
custom_hyphen_terms <- c("climate-change", "carbon-neutral", "fair-trade")
# 循环添加/修改词汇属性
for (term in custom_hyphen_terms) {
  lexeme <- vocab$get(term)
  if (is.null(lexeme)) {
    lexeme <- vocab$add(term)
  }
  # 标记为复合词,防止拆分
  lexeme$is_compound <- TRUE
}
# 再执行词形还原解析
test.sp <- spacy_parse(test.corpus, lemma = TRUE, entity = FALSE, pos = FALSE, tag = FALSE)

2. 绕过nounphrase_consolidate,手动处理连字符术语

放弃nounphrase_consolidate(),直接在spacy_parse的结果中针对性处理:

  • 识别原始文本中的连字符术语,对这部分直接保留(或替换为预先映射的 lemma 化完整术语),非连字符词汇正常使用 lemma。
# 执行解析时保留原始token和lemma
test.sp <- spacy_parse(test.corpus, lemma = TRUE, pos = TRUE, tag = FALSE, nounphrase = FALSE)
# 自定义处理逻辑:连字符术语保留原始形式,其他用lemma
test.sp$final_token <- ifelse(
  grepl("-", test.sp$token),
  test.sp$token,  # 若需要lemma化的连字符术语,可替换为预先定义的映射表值
  test.sp$lemma
)
# 转换为quanteda tokens
test.tokens <- as.tokens(test.sp, token = "final_token")
# 后续quanteda处理步骤
test.tokens <- tokens(test.tokens, remove_symbols = TRUE,
                      remove_numbers = TRUE,
                      remove_punct = TRUE,
                      remove_url = TRUE) %>% 
  tokens_tolower() %>% 
  tokens_select(pattern = stopwords("en"), selection = "remove")

3. 用quanteda的tokens_compound做后合并

如果spacyr已经拆分了连字符术语,可以用quanteda的复合词合并功能,将拆分后的词重新组合:

  • 提前定义目标复合词的词典,强制合并拆分后的词元。
# 定义需要合并的连字符术语词典(键为最终形式,值为拆分后的词)
hyphen_compounds <- dictionary(list(
  `climate-change` = c("climate", "change"),
  `carbon-neutral` = c("carbon", "neutral")
))
# 在quanteda分词后执行合并
test.tokens.3 <- test.tokens.3 %>% 
  tokens_compound(pattern = hyphen_compounds, concatenator = "-")

内容的提问来源于stack exchange,提问作者mgd_aus

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.19 22:05:12