R语言如何统计字符串向量内每条评论的总语句数与唯一语句数
R完全可以实现该需求,核心是按指定分隔符拆分字符串后做计数即可,以下是可直接运行的实现方案。
核心处理逻辑
- 以
//为固定分隔符拆分每条评论的字符串内容 - 清理拆分后单条语句首尾的多余空格,过滤空字符串,避免格式问题导致的计数误差
- 拆分得到的向量长度即为该条评论的总语句数
- 向量去重后的长度即为该条评论的唯一语句数
实现代码
基础R版本(无需安装额外依赖包)
# 读入示例数据 data<- structure(list(comment = c("Gary Goss made all notificiations and calls. Plow and vehicle // PD remains and MoDOT crews remain on scene. // 646 en route to check status of incident. // 646 was sent to check on a camera incident involving a tractor trailer of potatoes that needed to be uprighted from earlier in the night. ", "Crash with a TT and a sedan into the guarrdail. // 1 vehicle only sedan into the guard rail. // Guard Rail Damage located via camera. // MSHP C 160122961 TROOPER LACEY // NEED TO ESTIMATE " )), class = "data.frame", row.names = c(NA, -2L)) # 定义单条评论的计数函数 calc_sent_count <- function(comment_text) { # 固定分隔符拆分,清理首尾空格 sentences <- trimws(unlist(strsplit(comment_text, "//", fixed = TRUE))) # 过滤拆分产生的空字符串 sentences <- sentences[nzchar(sentences)] c( total_sentences = length(sentences), unique_sentences = length(unique(sentences)) ) } # 批量处理所有评论,将结果合并到原数据框 count_result <- t(sapply(data$comment, calc_sent_count)) data <- cbind(data, count_result)
运行后示例数据的输出结果为:第一条评论总语句数4、唯一语句数4;第二条评论总语句数5、唯一语句数5,和实际内容匹配。
Tidyverse版本(适配日常数据处理工作流)
如果日常使用dplyr+stringr处理数据,可以用更简洁的管道写法:
library(dplyr) library(stringr) data <- data %>% rowwise() %>% mutate( sent_split = list(trimws(str_split(comment, fixed("//"))[[1]])), sent_split = list(sent_split[nzchar(sent_split)]), total_sentences = length(sent_split), unique_sentences = length(unique(sent_split)) ) %>% select(-sent_split) %>% ungroup()
注意:上述代码添加了空格清理、空串过滤步骤,如果你的数据不存在分隔符前后多余空格、首尾多余分隔符的格式问题,可以删除这部分逻辑,直接对拆分结果计数即可。
内容的提问来源于stack exchange,提问作者Mustafa Kamal
相关产品推荐
相关产品推荐

