You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何移除tidyverse环境下字符型变量Class列的数值后缀

问题描述

已将数据从宽格式转换为长格式,但需要删除字符型变量Class中的数值(例如Diatom1、Diatom2这类条目,约120个),使用tidyverse包处理。现有代码如下:

#load raw data
here("data_raw", "Phytoplankton_Taxonomy_2022_clean_v1.csv")
PhytoRaw <- read_xlsx(here("data_raw","Phytoplankton_Taxonomy_2022_clean_v1.xlsx"))


#Make new column names, first merge the first two rows then pivot
data_head <- read_xlsx(here("data_raw","Phytoplankton_Taxonomy_2022_clean_v1.xlsx"), n_max = 2,col_names = FALSE)

new_names <- data_head %>% 
summarise(across(.fns = paste, collapse = "_")) %>%
unlist() %>% unname()

new_names

data <- read_xlsx(here("data_raw","Phytoplankton_Taxonomy_2022_clean_v1.xlsx"),
skip = 2, col_names = new_names)               

#convert data frame from a "wide" format to a "long" format
Phytolong <- data%>% pivot_longer(cols= !c("Date_NA", "Lake_NA", "Depth_NA"),
names_to= c("Class", "Species"),
names_sep = "_",
values_to= "Biomass")

#Rename Columns
Phytolong<- Phytolong %>% rename("Date" = "Date_NA", "Lake" = "Lake_NA", "Depth" = "Depth_NA")

尝试过mutate(Class = as.character(Class)),但该方法仅转换变量类型,无法移除字符中的数字,没有效果。

解决方案

要移除Class列中的数字,可使用tidyverse集成的stringr包处理字符串,结合mutate()调用str_remove_all()函数:

# 加载tidyverse(已包含stringr)
library(tidyverse)

# 移除Class列中所有数字字符
Phytolong <- Phytolong %>%
  mutate(Class = str_remove_all(Class, "\\d+"))

说明

  • \\d+是正则表达式,用于匹配一个或多个数字字符;
  • str_remove_all()会将Class列中所有匹配到的数字全部删除,比如Diatom1会变成Diatom,Chlorophyte3会变成Chlorophyte;
  • 如果仅需删除字符串末尾的数字(避免误删中间可能存在的数字),可以改用str_replace():
    Phytolong <- Phytolong %>%
      mutate(Class = str_replace(Class, "\\d+$", ""))
    
    其中$表示匹配字符串的末尾位置,仅删除末尾的数字序列。

内容的提问来源于stack exchange,提问作者Daniela

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.18 11:06:01