You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

为何quanteda的fcm()函数会显示零频率结果?

quanteda中fcm()零频率异常的原因解析

问题场景

你用quanteda处理就职演说语料库,生成特征共现矩阵(FCM)后,发现明明存在共现的词对(比如"since"和"and"),矩阵里却显示频率为0:

执行代码

library(dplyr)
library(tidyverse)
library(quanteda)

data_inaug <- data_corpus_inaugural

fcm_inaug <-  data_inaug %>% 
  tokens(remove_punct = TRUE,
         remove_symbols = TRUE,
         remove_numbers = TRUE,
         remove_url = TRUE) %>% 
  tokens_tolower() %>%
  fcm(context = "window", window = 5, count = "frequency")

# 查看子集
fcm_inaug[c("as", "for", "since", "because"), names(topfeatures(fcm_inaug))]

输出结果

Feature co-occurrence matrix of: 4 by 10 features.
         features
features  the our  we in and to is a  be for
  as        0 153 153  0   0  0 92 0 145  93
  for       0 210 165  0   0  0  0 0   0 110
  since     0   0   5  0   0  0  0 0   0   0
  because   0   0   0  0   0  0  0 0   0   0

但实际语料中存在大量"since"与"and"的共现,比如:

is noted to prove that <> truth and reason have maintained
rather more than forty-four years <> we declared our independence and
maxim of our policy ever <> the days of washington and
...

核心原因

问题出在fcm()的窗口方向设置上:

  • 当context = "window"时,window = 5这个参数默认只统计目标词右侧5个词的共现,完全不包含左侧的词。
  • 你举的例子里,大部分"and"都出现在"since"的左侧,这些共现根本没被纳入统计范围;而"since"右侧5个词内出现"and"的情况极少,所以矩阵里显示为0。

解决方法

如果需要统计目标词左右两侧的共现,把window参数改成长度为2的向量,明确指定左右窗口大小:

fcm_inaug <-  data_inaug %>% 
  tokens(remove_punct = TRUE,
         remove_symbols = TRUE,
         remove_numbers = TRUE,
         remove_url = TRUE) %>% 
  tokens_tolower() %>%
  # 设置左右各5个词的窗口
  fcm(context = "window", window = c(5,5), count = "frequency")

重新运行后,就能看到"since"与"and"的共现次数被正确统计了。

内容的提问来源于stack exchange,提问作者Znusgy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.29 20:00:19