如何在R语言中定位首条长度≥5的全大写单词的位置
在R语言中定位首条长度≥5的全大写单词位置及内容
需求:从给定字符向量的每个字符串里,找出第一个长度不小于5个字符的全大写单词,同时返回该单词的位置和单词本身。
示例数据
myvec <- c("FILT Words Here before CAPITALS words here after", "Words OUT ALLCAPS words MORECAPS words after", "Mywords PRE Before CAPLET more words after the capital Letters CAPLETTERS", "PRE CAP letters SPLIT here not before")
实现代码
# 把每个字符串拆分成单词列表 word_list <- strsplit(myvec, "\\s+") # 定义匹配规则:全大写字母且长度≥5 pattern <- "^[A-Z]{5,}$" # 逐个处理每组单词,找到第一个符合条件的单词及其位置 result <- lapply(word_list, function(words) { # 找出所有匹配的单词索引,取第一个 match_pos <- which(grepl(pattern, words))[1] list(pos = match_pos, word = ifelse(is.na(match_pos), NA, words[match_pos])) }) # 提取结果向量 desired_first_word_over4char_caps <- sapply(result, function(x) x$pos) desired_first_word_over4char <- sapply(result, function(x) x$word)
输出结果
> desired_first_word_over4char_caps [1] 5 3 4 4 > desired_first_word_over4char [1] "CAPITALS" "ALLCAPS" "CAPLET" "SPLIT"
关键说明
strsplit(myvec, "\\s+"):按任意数量的空白字符拆分字符串,得到每个字符串的单词列表^[A-Z]{5,}$:正则表达式确保匹配的单词全部由大写字母组成,且长度至少为5,^和$限制匹配范围是整个单词which(grepl(pattern, words))[1]:筛选出符合规则的单词索引,[1]直接取第一个匹配项,确保是首条符合条件的单词
内容的提问来源于stack exchange,提问作者Neal Barsch
相关产品推荐
相关产品推荐

