在R语言中创建支持多参数的文本缩略自定义函数的实现方法
R文本缩略功能自定义函数封装方案
前置依赖
处理字符串分割需要用到stringr包,运行前先加载:
# 未安装先执行安装 # install.packages("stringr") library(stringr)
封装完成的可调用函数
text_abbrev <- function(text, custom_stopwords = c("the", "really", "truly", "very"), min_abbrev_len = 5, extract_first_sent = TRUE) { # 1. 提取每个段落首句(按需开启) if (extract_first_sent) { text <- unlist(lapply(text, function(x) str_split(x, "\\.", simplify = T)[1])) } # 2. 移除指定停用词 text <- unlist(lapply(text, function(t) { words <- unlist(strsplit(t, " ")) # 过滤空字符串和停用词,统一转小写匹配避免大小写误差 words <- words[words != "" & !tolower(words) %in% tolower(custom_stopwords)] paste(words, collapse = " ") })) # 3. 长单词缩写,超出指定长度的部分替换为. regex_pattern <- paste0("(?<=\\w{", min_abbrev_len, "})\\w+") text <- gsub(regex_pattern, ".", text, perl = TRUE) # 4. 连续空格去重、清理首尾空格 text <- gsub("^ *|(?<= ) | *$", "", text, perl = TRUE) return(text) }
参数说明
text: 输入待处理文本,支持单条文本或多段落文本向量custom_stopwords: 自定义停用词向量,默认内置通用程度副词停用词min_abbrev_len: 触发缩写的单词最小长度,默认值为5,即长度超过5的单词保留前5位后用.缩写extract_first_sent: 逻辑值,是否开启段落首句提取功能,默认开启
使用示例
# 测试文本 test_input <- c( "The really great advantage of R language is that it has very rich text processing packages. This part of the content will be removed after extracting the first sentence.", "This is a second paragraph with truly supercalifragilisticexpialidocious long words. The following content will not be retained." ) # 调用函数 result <- text_abbrev(test_input) # 查看输出 print(result)
输出结果参考:
[1] "great advant. of R langua. is that it has rich text proces. packa." [2] "This is a second paragr. with super. long words"
优化说明
相对原有零散逻辑做了两处兼容优化:
- 停用词匹配改为不区分大小写,避免输入文本大小写不统一导致的停用词漏删
- 各功能可通过参数灵活开关,不需要的模块可直接关闭无需修改函数内部代码
内容的提问来源于stack exchange,提问作者Imran Luqman
相关产品推荐
相关产品推荐

