如何在R语言中将bcmort数据集的cohort变量拆分为Period和Area两个新变量
Hey there! Let's figure out how to split the cohort variable in the bcmort dataset (from the ISwR package) into two new variables: Period and Area. Since we know exactly the four levels of cohort, we have a few reliable ways to do this in R Studio.
Method 1: Tidyverse (dplyr + tidyr) – Intuitive and Concise
If you're already using the tidyverse ecosystem, this approach is super straightforward. We'll use string detection to pull out the key terms ("Historical" for period, "National" for area):
# Load required packages library(ISwR) library(tidyverse) # Load the dataset data(bcmort) # Split cohort into Period and Area bcmort <- bcmort %>% mutate( # Assign Period: "Historical" if the cohort starts with that term, else "Current" Period = case_when( str_detect(cohort, "^Historical") ~ "Historical", TRUE ~ "Current" ), # Assign Area: "National" if the cohort includes that term, else "Study" Area = case_when( str_detect(cohort, "National") ~ "National", TRUE ~ "Study" ) )
Method 2: Base R – No Extra Packages Needed
If you prefer sticking to base R, you can split the cohort strings and check for the key words directly:
library(ISwR) data(bcmort) # Convert cohort to character and split into individual words cohort_split <- strsplit(as.character(bcmort$cohort), " ") # Extract Period by checking for "Historical" in the split parts bcmort$Period <- sapply(cohort_split, function(parts) { if ("Historical" %in% parts) "Historical" else "Current" }) # Extract Area by checking for "National" in the split parts bcmort$Area <- sapply(cohort_split, function(parts) { if ("National" %in% parts) "National" else "Study" })
Method 3: Manual Factor Mapping – Precise and Reliable
Since cohort is a factor with fixed levels, we can create direct mappings for each level. This avoids any string matching edge cases:
library(ISwR) data(bcmort) # Create a mapping for Period period_mapping <- c( "Study Group" = "Current", "National study group" = "Current", "Historical Study Group" = "Historical", "Historical National Study Group" = "Historical" ) # Create a mapping for Area area_mapping <- c( "Study Group" = "Study", "National study group" = "National", "Historical Study Group" = "Study", "Historical National Study Group" = "National" ) # Apply mappings to create new variables bcmort$Period <- period_mapping[as.character(bcmort$cohort)] bcmort$Area <- area_mapping[as.character(bcmort$cohort)] # Optional: Convert new variables to factors (if you want categorical data) bcmort$Period <- as.factor(bcmort$Period) bcmort$Area <- as.factor(bcmort$Area)
Verify the Result
After running any of these methods, you can double-check the new variables with:
head(bcmort)
This will show you the first few rows of the dataset, including your new Period and Area columns.
内容的提问来源于stack exchange,提问作者Maximus

