遍历字符串列表统计词频遇strsplit错误,求R语言解决方案
R语言词频统计报错解决:non-character argument
问题重现
你创建了包含两个字符串的列表:
index <- lst(string1="there is guy is there there", string2="have a have a ball a")
编写了单个字符串可用的词频统计函数:
get_wordcount <- function(text) { wordcount <- sort(table(unlist(strsplit(text, " "))), decreasing = TRUE) return(wordcount) }
但用for循环遍历列表时触发错误:
for (n in 1:length(index)) { assign(paste(names(index[n]), "wordcount", sep = "_"), get_wordcount(index[n])) }
错误信息:
Error in strsplit(text, " ") : non-character argument
错误根源
index[n]提取的是长度为1的子列表,不是字符串本身。比如index[1]返回的是带命名的列表结构,而非单纯的字符向量,而strsplit仅接受字符类型输入,因此报错。
修复方案
把循环里的index[n]改成index[[n]](双中括号提取列表元素的实际值),同时优化命名获取逻辑:
for (n in 1:length(index)) { assign(paste(names(index)[n], "wordcount", sep = "_"), get_wordcount(index[[n]])) }
更推荐用lapply批量处理,结果统一存在列表中,避免生成零散的全局变量:
# 批量处理所有字符串 wordcount_list <- lapply(index, get_wordcount) # 若需要把列表元素转为全局变量,用list2env list2env(wordcount_list, envir = .GlobalEnv)
运行结果
修复后即可得到期望的词频统计:
> string1_wordcount there is guy 3 2 1 > string2_wordcount a have ball 3 2 1
内容的提问来源于stack exchange,提问作者IL DOGE
相关产品推荐
相关产品推荐

