使用quanteda.textstats时遇‘undefined columns selected’错误求助
解决quanteda中tokens_compound调用的"undefined columns selected"错误
问题场景
在主题模型项目中,使用quanteda与quanteda.textstats构建二元语法后,调用tokens_compound函数时出现以下错误:
Error in `[.data.frame`(col, col$z > 1) : undefined columns selected
已确认col数据框中存在z列且有值大于1,使用janeaustenr包的sensesensibility数据集可复现该问题,原始代码如下:
library(janeaustenr) library(tidyverse) library(quanteda) library(quanteda.textstats) data(sensesensibility) sensesensibility <- as.data.frame(sensesensibility) ##add id column to the data frame sensesensibility <- sensesensibility %>% mutate(id = row_number()) %>% slice(13:30) ##first few lines aren't full lines of text corpus <- corpus(sensesensibility, docid_field = "id", text_field = "sensesensibility") tokens <- corpus %>% tokens(remove_punct = TRUE, remove_numbers = TRUE, remove_symbols = TRUE, remove_separators = TRUE) %>% tokens_tolower() ##see if it worked (it did) tokens col <- textstat_collocations(tokens, size = 2) tokens.col <- tokens_compound(tokens, pattern = col[col$z > 1], concatenator = "_")
错误原因
核心问题在于数据框索引的误用:
col是textstat_collocations返回的collocations类数据框,col$z > 1是一个与行数等长的逻辑向量。- 在R中,
data.frame的索引格式为df[行索引, 列索引],若省略逗号只写df[逻辑向量],R会默认用该向量筛选列而非行。由于逻辑向量长度(行数)远大于列数,直接引发"undefined columns selected"错误。 - 另外,
tokens_compound的pattern参数需要接收collocations对象或有效词元模式,必须确保传入的是符合条件的行数据,而非错误筛选的列。
解决方案
补全行索引的逗号,明确筛选符合条件的行,有两种简洁写法:
- 使用完整行索引语法:
tokens.col <- tokens_compound(tokens, pattern = col[col$z > 1, ], concatenator = "_")
- 使用
subset函数更清晰地筛选行:
tokens.col <- tokens_compound(tokens, pattern = subset(col, z > 1), concatenator = "_")
修正后的完整代码
library(janeaustenr) library(tidyverse) library(quanteda) library(quanteda.textstats) data(sensesensibility) sensesensibility <- as.data.frame(sensesensibility) ##add id column to the data frame sensesensibility <- sensesensibility %>% mutate(id = row_number()) %>% slice(13:30) ##first few lines aren't full lines of text corpus <- corpus(sensesensibility, docid_field = "id", text_field = "sensesensibility") tokens <- corpus %>% tokens(remove_punct = TRUE, remove_numbers = TRUE, remove_symbols = TRUE, remove_separators = TRUE) %>% tokens_tolower() col <- textstat_collocations(tokens, size = 2) # 修正后的调用方式 tokens.col <- tokens_compound(tokens, pattern = col[col$z > 1, ], concatenator = "_")
内容的提问来源于stack exchange,提问作者seder163
相关产品推荐
相关产品推荐

