R语言5个字符串处理实操问题咨询
R字符串处理5个需求实现方案
以下实现默认使用stringr包(tidyverse生态常用的字符串处理包),也同步提供base R兼容写法。
需求1:提取第三个与第四个点号之间的数字,去除前导0
实现代码
library(stringr) # stringr写法 extract1 <- function(x) { matched <- str_match(x, "^(?:[^.]+\\.){3}(\\d+)\\.")[,2] sub("^0+", "", matched) } # base R写法 extract1_base <- function(x) { matched <- regmatches(x, regexec("^(?:[^.]+\\.){3}(\\d+)\\.", x))[[1]][2] sub("^0+", "", matched) }
测试验证
输入"01.030.131.012962.000000 -1/1 -",输出结果为12962。
需求2:提取倒数第一个和倒数第二个点号之间的数字
实现代码
# stringr写法 extract2 <- function(x) { str_match(x, "\\.(\\d+)\\.[^.]*$")[,2] } # base R写法 extract2_base <- function(x) { regmatches(x, regexec("\\.(\\d+)\\.[^.]*$", x))[[1]][2] }
测试验证
输入"33.31.17811849.0",输出结果为17811849。
需求3:提取第二个与第三个斜杠之间的数字,去除前导0
实现代码
# stringr写法 extract3 <- function(x) { matched <- str_match(x, "^(?:[^/]+/){2}(\\d+)/")[,2] sub("^0+", "", matched) } # base R写法 extract3_base <- function(x) { matched <- regmatches(x, regexec("^(?:[^/]+/){2}(\\d+)/", x))[[1]][2] sub("^0+", "", matched) }
测试验证
输入"2001/96/00000036/0",输出结果为36。
需求4:提取第二个点号之后的所有数字,剔除内容左侧所有的0、点号、短横线
实现代码
# stringr写法 extract4 <- function(x) { # 截取第二个点之后的全部内容 after_dot <- sub("^(?:[^.]+\\.){2}", "", x) # 剔除开头连续的0、.、- trimed <- sub("^[0.-]+", "", after_dot) # 提取所有数字并拼接 str_replace_all(trimed, "\\D", "") } # base R写法 extract4_base <- function(x) { after_dot <- sub("^(?:[^.]+\\.){2}", "", x) trimed <- sub("^[0.-]+", "", after_dot) gsub("\\D", "", trimed) }
测试验证
输入"0118.11.54.779-9",输出结果为547799。
需求5:提取775、776或746三个数字序列之后的前7位数字
实现代码
# stringr写法 extract5 <- function(x) { str_match(x, "(?:775|776|746)(\\d{7})")[,2] } # base R写法 extract5_base <- function(x) { regmatches(x, regexec("(?:775|776|746)(\\d{7})", x))[[1]][2] }
测试验证
输入"0306920189907750240991000",输出结果为0240991。
内容的提问来源于stack exchange,提问作者Barbara Jacqueline
相关产品推荐
相关产品推荐

