如何用基础R统计字符串中长度≥3的单词数量?
错误原因分析
你遇到的问题核心是对strsplit()返回值类型的误解:
strsplit()函数返回的是列表对象,哪怕只处理单个字符串,结果也是包含一个向量元素的列表。- 你写的
for (word in words)循环中,每次迭代的word其实是整个单词向量,而非单个单词。此时nchar(word)会生成所有单词长度组成的向量,将这个向量传入if条件时,就触发了错误——基础R的if语句只能处理长度为1的逻辑判断,无法直接处理向量。
基础R解决方案
以下几种方式都不需要外部包,完全用基础R实现:
方式一:修正循环逻辑(和你Python思路对齐)
先从strsplit返回的列表中提取出单词向量,再遍历单个单词:
sentence <- 'I have a string sentence but i do not know how to get three lettered words from it' words <- strsplit(sentence, split = ' ')[[1]] # 用[[1]]提取列表内的向量 count <- 0 for (word in words) { if (nchar(word) >= 3) { count <- count + 1 } } print(count) # 输出12
方式二:向量化操作(最简洁)
利用R的向量化特性,直接计算符合条件的单词数量:
sentence <- 'I have a string sentence but i do not know how to get three lettered words from it' words <- strsplit(sentence, split = ' ')[[1]] count <- sum(nchar(words) >= 3) # 逻辑向量求和,TRUE=1,FALSE=0 print(count) # 输出12
方式三:类似Python的Filter实现
用基础R的Filter()函数筛选符合条件的单词,再统计长度:
sentence <- 'I have a string sentence but i do not know how to get three lettered words from it' words <- strsplit(sentence, split = ' ')[[1]] filtered <- Filter(function(w) nchar(w) >= 3, words) count <- length(filtered) print(count) # 输出12
内容的提问来源于stack exchange,提问作者Veki
相关产品推荐
相关产品推荐

