如何在R语言中将形如‘Male25’的列拆分为性别与年龄两列?
Hey there! Splitting a column formatted like Male25 (letters for gender, digits for age) into two separate columns is a super common data cleaning task in R. Here are the most efficient, readable methods to get this done:
Method 1: Using tidyr::separate_wider_regex (Tidyverse Approach)
If you're already working with the tidyverse (which most R users rely on for data manipulation), separate_wider_regex is the most intuitive way. It lets you explicitly define regex patterns for each new column:
# Load the tidyverse package (includes tidyr and dplyr) library(tidyverse) # Example data frame df <- tibble(combined = c("Male25", "Female32", "Male19", "Female45")) # Split the column df_split <- df %>% separate_wider_regex( combined, patterns = list( gender = "[A-Za-z]+", # Match all letters for gender age = "\\d+" # Match all digits for age ), convert = TRUE # Auto-convert age to numeric type ) # View the result df_split
This method cleansly splits the column and handles type conversion in one step—perfect for streamlined workflows.
Method 2: Base R (No External Packages)
If you prefer sticking to base R without loading extra libraries, use sub() to extract each component directly:
# Example data frame df <- data.frame(combined = c("Male25", "Female32", "Male19", "Female45")) # Extract gender (remove all digits from the string) df$gender <- sub("\\d+", "", df$combined) # Extract age (remove all letters, then convert to numeric) df$age <- as.numeric(sub("[A-Za-z]+", "", df$combined)) # Check the result df
This is lightweight and great for quick, one-off tasks where you don't want to pull in additional dependencies.
Method 3: Using stringr::str_extract (Tidyverse Alternative)
Another flexible tidyverse option is using stringr's str_extract to pull out matching patterns individually:
library(stringr) library(dplyr) df <- tibble(combined = c("Male25", "Female32", "Male19", "Female45")) df_split <- df %>% mutate( gender = str_extract(combined, "[A-Za-z]+"), age = as.numeric(str_extract(combined, "\\d+")) ) df_split
This approach gives you granular control over each extraction—handy if you need to adjust one column's logic without affecting the other.
Key Notes
- All these methods assume your data follows the letters first, digits second pattern. If you have edge cases (like mixed case or non-standard entries), tweak the regex (e.g., use
[A-Za-z]+to handle both upper and lowercase). - Always convert the
agecolumn to numeric withas.numeric()if you plan to do numerical analysis (like calculating averages or filtering by age ranges).
内容的提问来源于stack exchange,提问作者Neoleogeo

