如何移除tidyverse环境下字符型变量Class列的数值后缀
问题描述
已将数据从宽格式转换为长格式,但需要删除字符型变量Class中的数值(例如Diatom1、Diatom2这类条目,约120个),使用tidyverse包处理。现有代码如下:
#load raw data here("data_raw", "Phytoplankton_Taxonomy_2022_clean_v1.csv") PhytoRaw <- read_xlsx(here("data_raw","Phytoplankton_Taxonomy_2022_clean_v1.xlsx")) #Make new column names, first merge the first two rows then pivot data_head <- read_xlsx(here("data_raw","Phytoplankton_Taxonomy_2022_clean_v1.xlsx"), n_max = 2,col_names = FALSE) new_names <- data_head %>% summarise(across(.fns = paste, collapse = "_")) %>% unlist() %>% unname() new_names data <- read_xlsx(here("data_raw","Phytoplankton_Taxonomy_2022_clean_v1.xlsx"), skip = 2, col_names = new_names) #convert data frame from a "wide" format to a "long" format Phytolong <- data%>% pivot_longer(cols= !c("Date_NA", "Lake_NA", "Depth_NA"), names_to= c("Class", "Species"), names_sep = "_", values_to= "Biomass") #Rename Columns Phytolong<- Phytolong %>% rename("Date" = "Date_NA", "Lake" = "Lake_NA", "Depth" = "Depth_NA")
尝试过mutate(Class = as.character(Class)),但该方法仅转换变量类型,无法移除字符中的数字,没有效果。
解决方案
要移除Class列中的数字,可使用tidyverse集成的stringr包处理字符串,结合mutate()调用str_remove_all()函数:
# 加载tidyverse(已包含stringr) library(tidyverse) # 移除Class列中所有数字字符 Phytolong <- Phytolong %>% mutate(Class = str_remove_all(Class, "\\d+"))
说明
\\d+是正则表达式,用于匹配一个或多个数字字符;str_remove_all()会将Class列中所有匹配到的数字全部删除,比如Diatom1会变成Diatom,Chlorophyte3会变成Chlorophyte;- 如果仅需删除字符串末尾的数字(避免误删中间可能存在的数字),可以改用
str_replace():
其中Phytolong <- Phytolong %>% mutate(Class = str_replace(Class, "\\d+$", ""))$表示匹配字符串的末尾位置,仅删除末尾的数字序列。
内容的提问来源于stack exchange,提问作者Daniela
相关产品推荐
相关产品推荐

