R中如何用正则提取df字段中Type N字符串并转换为type_n格式
R语言Type格式字符串转换实现方案
方法1:tidyverse系列stringr包实现(推荐)
操作最简,可读性高,适合已经在用tidyverse生态的场景:
- 假设存储原始字符串的列名为
raw_col,生成的新列名为new_type - 先加载依赖包:
library(tidyverse) - 单管道链完成提取+转换:
df <- df %>% mutate( new_type = str_extract(raw_col, "Type\\s\\d") %>% # 仅提取Type+空格+数字的目标片段,原列有其他内容也不受影响 str_to_lower() %>% # 大写转小写 str_replace(" ", "_") # 空格替换为下划线 )
如果原列仅包含Type 1这类目标内容,没有其他冗余字符串,可以省略提取步骤,直接转换:
df <- df %>% mutate(new_type = str_to_lower(raw_col) %>% str_replace(" ", "_"))
方法2:R基础包实现(无需安装第三方依赖)
不需要额外安装包,直接用内置函数完成:
- 原列有冗余内容的场景:
df$new_type <- tolower(gsub(" ", "_", regmatches(df$raw_col, regexpr("Type\\s\\d", df$raw_col))))
- 原列无冗余内容的场景:
df$new_type <- tolower(gsub(" ", "_", df$raw_col))
内容的提问来源于stack exchange,提问作者MLE
相关产品推荐
相关产品推荐

