如何在R中对长字符串进行首尾完整单词截断并拼接?
R实现长文本前后截断并保留完整单词
你可以通过自定义函数结合sapply和strsplit来实现需求,核心思路是分别处理文本的开头和结尾部分,再用...拼接:
完整代码
# 构造示例数据框 df <- data.frame( text = c( "This is a very long sentence that I would like to trim because I might need to put it as a label somewhere", "This is another very long sentence that I would also like to trim because I might need to put it as who knows what" ), col2 = c("1234", "5678") ) # 定义处理单条文本的函数 trim_text <- function(words) { total_words <- length(words) if (total_words == 0) return("") # 截取开头20字符内的完整单词 start_cumsum <- cumsum(nchar(words)) start_idx <- max(which(start_cumsum <= 20)) # 截取结尾20字符内的完整单词(从后往前计算) reversed_words <- rev(words) end_cumsum <- cumsum(nchar(reversed_words)) end_idx <- max(which(end_cumsum <= 20)) end_start_pos <- total_words - end_idx + 1 # 处理文本过短的情况:如果开头和结尾重叠,直接返回原文本 if (start_idx >= end_start_pos) { return(paste(words, collapse = " ")) } # 拼接结果 start_part <- paste(words[1:start_idx], collapse = " ") end_part <- paste(words[end_start_pos:total_words], collapse = " ") paste(start_part, end_part, sep = "...") } # 应用函数生成新列 df$short_txt <- sapply(strsplit(df$text, " "), trim_text) # 查看结果 df$short_txt
运行结果
[1] "This is a very long...it as a label somewhere" [2] "This is another very...it as who knows what"
代码说明
- 开头截取:通过
cumsum(nchar(words))计算单词长度的累计和,找到累计和不超过20的最后一个单词,截取前面所有单词。 - 结尾截取:先反转单词列表,用同样的累计和方法找到末尾不超过20字符的单词,再反转回原顺序。
- 边界处理:如果文本总长度较短,开头和结尾的截取范围重叠时,直接返回原文本避免重复内容。
内容的提问来源于stack exchange,提问作者Unai Vicente
相关产品推荐
相关产品推荐

