R语言如何使用stringr的word函数提取句子至多前3个单词
解决方法
方法1:结合单词计数动态调用word()函数
先统计每个句子的实际单词数,调用word()时结束位置取实际单词数和3的较小值,避免返回NA。
library(stringr) sentences <- c("Jane saw a cat", "Jane sat down", "Jane sat", "Jane") # 统计每个句子的单词数 word_cnt <- str_count(sentences, "\\S+") # 动态取前min(3, 单词数)个词 res <- mapply(function(s, n) word(s, 1, min(3, n)), sentences, word_cnt, USE.NAMES = F)
运行后res的输出为:"Jane saw a" "Jane sat down" "Jane sat" "Jane"
方法2:正则表达式直接替换(更简洁)
用正则匹配句子开头至多3个单词,直接删掉后续内容,无需循环和额外计数:
library(stringr) sentences <- c("Jane saw a cat", "Jane sat down", "Jane sat", "Jane") res <- str_replace(sentences, "^((\\S+\\s+){0,2}\\S+).*$", "\\1")
方法3:纯Base R实现(无需加载第三方包)
用基础函数拆分字符串后取前3项再拼接:
sentences <- c("Jane saw a cat", "Jane sat down", "Jane sat", "Jane") sent_list <- strsplit(sentences, "\\s+") res <- sapply(sent_list, function(x) paste(head(x, 3), collapse = " "))
内容的提问来源于stack exchange,提问作者dano_
相关产品推荐
相关产品推荐

