You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用quanteda.textstats时遇‘undefined columns selected’错误求助

解决quanteda中tokens_compound调用的"undefined columns selected"错误

问题场景

在主题模型项目中,使用quanteda与quanteda.textstats构建二元语法后,调用tokens_compound函数时出现以下错误:

Error in `[.data.frame`(col, col$z > 1) : undefined columns selected

已确认col数据框中存在z列且有值大于1,使用janeaustenr包的sensesensibility数据集可复现该问题,原始代码如下:

library(janeaustenr)
library(tidyverse)
library(quanteda)
library(quanteda.textstats)

data(sensesensibility)

sensesensibility <- as.data.frame(sensesensibility)

##add id column to the data frame
sensesensibility <- sensesensibility %>% 
  mutate(id = row_number()) %>% 
  slice(13:30) ##first few lines aren't full lines of text

corpus <- corpus(sensesensibility,
                     docid_field = "id",
                     text_field = "sensesensibility")


tokens <- corpus %>%
  tokens(remove_punct = TRUE, 
         remove_numbers = TRUE, 
         remove_symbols = TRUE, 
         remove_separators = TRUE) %>%
  tokens_tolower()

##see if it worked (it did)
tokens

col <- textstat_collocations(tokens, size = 2)

tokens.col <- tokens_compound(tokens, pattern = col[col$z > 1], concatenator = "_")

错误原因

核心问题在于数据框索引的误用:

  • col是textstat_collocations返回的collocations类数据框,col$z > 1是一个与行数等长的逻辑向量。
  • 在R中,data.frame的索引格式为df[行索引, 列索引],若省略逗号只写df[逻辑向量],R会默认用该向量筛选列而非行。由于逻辑向量长度(行数)远大于列数,直接引发"undefined columns selected"错误。
  • 另外,tokens_compound的pattern参数需要接收collocations对象或有效词元模式,必须确保传入的是符合条件的行数据,而非错误筛选的列。

解决方案

补全行索引的逗号,明确筛选符合条件的行,有两种简洁写法:

  1. 使用完整行索引语法:
tokens.col <- tokens_compound(tokens, pattern = col[col$z > 1, ], concatenator = "_")
  1. 使用subset函数更清晰地筛选行:
tokens.col <- tokens_compound(tokens, pattern = subset(col, z > 1), concatenator = "_")

修正后的完整代码

library(janeaustenr)
library(tidyverse)
library(quanteda)
library(quanteda.textstats)

data(sensesensibility)

sensesensibility <- as.data.frame(sensesensibility)

##add id column to the data frame
sensesensibility <- sensesensibility %>% 
  mutate(id = row_number()) %>% 
  slice(13:30) ##first few lines aren't full lines of text

corpus <- corpus(sensesensibility,
                     docid_field = "id",
                     text_field = "sensesensibility")


tokens <- corpus %>%
  tokens(remove_punct = TRUE, 
         remove_numbers = TRUE, 
         remove_symbols = TRUE, 
         remove_separators = TRUE) %>%
  tokens_tolower()

col <- textstat_collocations(tokens, size = 2)

# 修正后的调用方式
tokens.col <- tokens_compound(tokens, pattern = col[col$z > 1, ], concatenator = "_")

内容的提问来源于stack exchange,提问作者seder163

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.25 06:03:24