如何在R语言中统计字符串中的非空片段数量?
嘿,我来帮你搞定这个问题!先把你的需求和遇到的卡点理清楚:
你的问题场景
你有这么一段字符串(注意原字符串里的\"是R自动打印的转义符,实际不存在):
"Jenna and Alex were making cupcakes.", "Jenna asked Alex whether all were ready to be frosted.", "Alex said that", " some of them ", "were.", "He added", "that", "the rest", "would be", "ready", "soon.", ""
你想要统计其中非空片段的数量,正确结果应该是11,但直接把这段字符串转成向量时,R会把它当成单个元素,向量长度始终是1,完全达不到统计目的。
核心原因
你手里的是一个包含多个双引号包裹片段的单字符串,R不会自动识别引号里的内容为独立元素,必须先把这些片段提取出来,再过滤空值统计数量。
两种可行的解决方法
方法1:用stringr包(推荐,语法更简洁)
stringr是tidyverse生态里的字符串处理工具包,str_match_all()可以精准提取引号内的内容(包括捕获组),步骤如下:
# 定义原始字符串(去掉R自动添加的转义符) raw_str <- '"Jenna and Alex were making cupcakes.", "Jenna asked Alex whether all were ready to be frosted.", "Alex said that", " some of them ", "were.", "He added", "that", "the rest", "would be", "ready", "soon.", ""' # 先安装包(如果没装过的话) # install.packages("stringr") library(stringr) # 提取所有双引号包裹的内容(只取引号内的部分,不含引号) extracted_fragments <- str_match_all(raw_str, '"([^"]*)"')[[1]][,2] # 过滤空字符串,统计非空片段数量 non_empty_count <- sum(nzchar(extracted_fragments)) print(non_empty_count) # 输出11,完全符合预期
这里的正则表达式"([^"]*)"的意思是:匹配双引号,然后捕获所有不是双引号的内容,直到下一个双引号,这样就能精准提取每个片段。
方法2:基础R实现(无需额外安装包)
如果你不想装新包,用基础R的gregexpr+regmatches组合也能实现:
raw_str <- '"Jenna and Alex were making cupcakes.", "Jenna asked Alex whether all were ready to be frosted.", "Alex said that", " some of them ", "were.", "He added", "that", "the rest", "would be", "ready", "soon.", ""' # 找到所有匹配的位置并提取内容 matched_content <- regmatches(raw_str, gregexpr('"([^"]*)"', raw_str))[[1]] # 去掉每个片段前后的双引号 extracted_fragments <- sub('^"|"$', '', matched_content) # 统计非空片段数量 non_empty_count <- sum(nzchar(extracted_fragments)) print(non_empty_count) # 输出11
小提示
nzchar()函数用来判断字符串是否非空,比extracted_fragments != ""更稳妥,因为它能处理一些特殊的空白字符情况。
内容的提问来源于stack exchange,提问作者Taranee Cao

