为何quanteda的fcm()函数会显示零频率结果?
quanteda中fcm()零频率异常的原因解析
问题场景
你用quanteda处理就职演说语料库,生成特征共现矩阵(FCM)后,发现明明存在共现的词对(比如"since"和"and"),矩阵里却显示频率为0:
执行代码
library(dplyr) library(tidyverse) library(quanteda) data_inaug <- data_corpus_inaugural fcm_inaug <- data_inaug %>% tokens(remove_punct = TRUE, remove_symbols = TRUE, remove_numbers = TRUE, remove_url = TRUE) %>% tokens_tolower() %>% fcm(context = "window", window = 5, count = "frequency") # 查看子集 fcm_inaug[c("as", "for", "since", "because"), names(topfeatures(fcm_inaug))]
输出结果
Feature co-occurrence matrix of: 4 by 10 features. features features the our we in and to is a be for as 0 153 153 0 0 0 92 0 145 93 for 0 210 165 0 0 0 0 0 0 110 since 0 0 5 0 0 0 0 0 0 0 because 0 0 0 0 0 0 0 0 0 0
但实际语料中存在大量"since"与"and"的共现,比如:
is noted to prove that <
> truth and reason have maintained
rather more than forty-four years <> we declared our independence and
maxim of our policy ever <> the days of washington and
...
核心原因
问题出在fcm()的窗口方向设置上:
- 当
context = "window"时,window = 5这个参数默认只统计目标词右侧5个词的共现,完全不包含左侧的词。 - 你举的例子里,大部分"and"都出现在"since"的左侧,这些共现根本没被纳入统计范围;而"since"右侧5个词内出现"and"的情况极少,所以矩阵里显示为0。
解决方法
如果需要统计目标词左右两侧的共现,把window参数改成长度为2的向量,明确指定左右窗口大小:
fcm_inaug <- data_inaug %>% tokens(remove_punct = TRUE, remove_symbols = TRUE, remove_numbers = TRUE, remove_url = TRUE) %>% tokens_tolower() %>% # 设置左右各5个词的窗口 fcm(context = "window", window = c(5,5), count = "frequency")
重新运行后,就能看到"since"与"and"的共现次数被正确统计了。
内容的提问来源于stack exchange,提问作者Znusgy
相关产品推荐
相关产品推荐

