You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言:批量重命名数据框中列的字符变量

Efficient Bulk Mapping for Categorical Column with 1M Rows

First, let's tackle this with a priority-based mapping approach—this lets you avoid writing separate code for each of the 50 unique values, and it's optimized for large datasets (1M rows is totally manageable here).

Step 1: Define Your Mapping Rules

First, create a named vector that maps your 10 target categories to their matching keywords, ordered by priority. For example, based on your examples:

  • Any string containing "Special Needs" should map to "Special Needs" (highest priority)
  • Strings with "Math & Science" map to "Math"
  • And so on for your remaining target categories.

Here's how to structure this:

# Define priority mapping: target category -> matching keyword
category_mapping <- c(
  "Special Needs" = "Special Needs",
  "Math" = "Math & Science",
  "Literacy & Language" = "Literacy & Language",
  "History & Civics" = "History & Civics",
  "Health & Sports" = "Health & Sports",
  "Music & The Arts" = "Music & The Arts",
  "Applied Learning" = "Applied Learning"
  # Add your remaining 3 target categories here
)

Note: Order matters here! Earlier entries have higher priority. So if a string has both "Applied Learning" and "Special Needs", it will pick "Special Needs" because it's listed first.

Step 2: Bulk Map the Column (Efficient Methods)

We'll use vectorized operations (no slow loops!) to handle the 1M rows quickly. Below are two popular options:

Option 1: Using dplyr & stringr (Readable, Good for Most Cases)

If you're already using the tidyverse, this is straightforward:

library(dplyr)
library(stringr)

projectdf <- projectdf %>%
  mutate(
    ProjectSubject_Clean = case_when(
      # Match each rule in priority order
      str_detect(ProjectSubject, category_mapping["Special Needs"]) ~ names(category_mapping["Special Needs"]),
      str_detect(ProjectSubject, category_mapping["Math"]) ~ names(category_mapping["Math"]),
      str_detect(ProjectSubject, category_mapping["Literacy & Language"]) ~ names(category_mapping["Literacy & Language"]),
      # Add the rest of your 10 target category rules here...
      # Fallback for edge cases (adjust if you don't need this)
      TRUE ~ "Other"
    )
  )

Option 2: Using data.table (Ultra-Fast for Large Datasets)

For 1M rows, data.table will outperform dplyr in speed. Here's how to implement it:

library(data.table)

# Convert to data.table for optimized performance
setDT(projectdf)

# Create the cleaned column using vectorized fcase
projectdf[, ProjectSubject_Clean := fcase(
  str_detect(ProjectSubject, category_mapping["Special Needs"]), names(category_mapping["Special Needs"]),
  str_detect(ProjectSubject, category_mapping["Math"]), names(category_mapping["Math"]),
  str_detect(ProjectSubject, category_mapping["Literacy & Language"]), names(category_mapping["Literacy & Language"]),
  # Add remaining target category rules here...
  default = "Other"
)]

Step 3: Verify Your Mapping

Before applying to the full dataset, test with your unique values to ensure everything maps correctly:

# Generate a lookup table for unique original values and their cleaned versions
unique_mapping <- data.frame(
  Original = unique(projectdf$ProjectSubject),
  Clean = sapply(unique(projectdf$ProjectSubject), function(x) {
    # Find the first matching priority rule
    match_idx <- which(str_detect(x, category_mapping))[1]
    ifelse(is.na(match_idx), "Other", names(category_mapping)[match_idx])
  })
)

# Review the results to adjust priorities or mappings if needed
print(unique_mapping)

This lets you confirm all 50 unique values are correctly mapped to your 10 target categories, and tweak the priority order if any edge cases pop up.

Why This Works

  • No manual per-value code: You only define your 10 target categories once, and the logic handles all 50 unique strings automatically.
  • Speed: Vectorized operations like case_when and fcase process 1M rows in seconds, no slow loops required.
  • Flexibility: Adjust priorities, add new target categories, or modify keywords by just updating the category_mapping vector.

内容的提问来源于stack exchange,提问作者Marcel

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.29 08:35:43