R语言:批量重命名数据框中列的字符变量
First, let's tackle this with a priority-based mapping approach—this lets you avoid writing separate code for each of the 50 unique values, and it's optimized for large datasets (1M rows is totally manageable here).
Step 1: Define Your Mapping Rules
First, create a named vector that maps your 10 target categories to their matching keywords, ordered by priority. For example, based on your examples:
- Any string containing "Special Needs" should map to "Special Needs" (highest priority)
- Strings with "Math & Science" map to "Math"
- And so on for your remaining target categories.
Here's how to structure this:
# Define priority mapping: target category -> matching keyword category_mapping <- c( "Special Needs" = "Special Needs", "Math" = "Math & Science", "Literacy & Language" = "Literacy & Language", "History & Civics" = "History & Civics", "Health & Sports" = "Health & Sports", "Music & The Arts" = "Music & The Arts", "Applied Learning" = "Applied Learning" # Add your remaining 3 target categories here )
Note: Order matters here! Earlier entries have higher priority. So if a string has both "Applied Learning" and "Special Needs", it will pick "Special Needs" because it's listed first.
Step 2: Bulk Map the Column (Efficient Methods)
We'll use vectorized operations (no slow loops!) to handle the 1M rows quickly. Below are two popular options:
Option 1: Using dplyr & stringr (Readable, Good for Most Cases)
If you're already using the tidyverse, this is straightforward:
library(dplyr) library(stringr) projectdf <- projectdf %>% mutate( ProjectSubject_Clean = case_when( # Match each rule in priority order str_detect(ProjectSubject, category_mapping["Special Needs"]) ~ names(category_mapping["Special Needs"]), str_detect(ProjectSubject, category_mapping["Math"]) ~ names(category_mapping["Math"]), str_detect(ProjectSubject, category_mapping["Literacy & Language"]) ~ names(category_mapping["Literacy & Language"]), # Add the rest of your 10 target category rules here... # Fallback for edge cases (adjust if you don't need this) TRUE ~ "Other" ) )
Option 2: Using data.table (Ultra-Fast for Large Datasets)
For 1M rows, data.table will outperform dplyr in speed. Here's how to implement it:
library(data.table) # Convert to data.table for optimized performance setDT(projectdf) # Create the cleaned column using vectorized fcase projectdf[, ProjectSubject_Clean := fcase( str_detect(ProjectSubject, category_mapping["Special Needs"]), names(category_mapping["Special Needs"]), str_detect(ProjectSubject, category_mapping["Math"]), names(category_mapping["Math"]), str_detect(ProjectSubject, category_mapping["Literacy & Language"]), names(category_mapping["Literacy & Language"]), # Add remaining target category rules here... default = "Other" )]
Step 3: Verify Your Mapping
Before applying to the full dataset, test with your unique values to ensure everything maps correctly:
# Generate a lookup table for unique original values and their cleaned versions unique_mapping <- data.frame( Original = unique(projectdf$ProjectSubject), Clean = sapply(unique(projectdf$ProjectSubject), function(x) { # Find the first matching priority rule match_idx <- which(str_detect(x, category_mapping))[1] ifelse(is.na(match_idx), "Other", names(category_mapping)[match_idx]) }) ) # Review the results to adjust priorities or mappings if needed print(unique_mapping)
This lets you confirm all 50 unique values are correctly mapped to your 10 target categories, and tweak the priority order if any edge cases pop up.
Why This Works
- No manual per-value code: You only define your 10 target categories once, and the logic handles all 50 unique strings automatically.
- Speed: Vectorized operations like
case_whenandfcaseprocess 1M rows in seconds, no slow loops required. - Flexibility: Adjust priorities, add new target categories, or modify keywords by just updating the
category_mappingvector.
内容的提问来源于stack exchange,提问作者Marcel

