如何拆分列唯一值为新列并移除原列?含mtcars实操场景
Hey there! Let's tackle your two R data frame questions with practical, commonly-used solutions:
This is typically referred to as one-hot encoding (creating indicator variables for each category). Here are two reliable approaches:
Using tidyverse (dplyr + tidyr)
First, load the tidyverse package if you haven't already. We'll use pivot_wider() to reshape the data, plus a helper column to mark presence of each category:
library(tidyverse) # Example data frame to demonstrate df <- tibble( id = 1:5, category = c("A", "B", "A", "C", "B") ) # Process the data df_transformed <- df %>% mutate(indicator = 1) %>% # Create a helper column with value 1 pivot_wider( names_from = category, # Column to split into new columns values_from = indicator, # Use the helper column for values values_fill = 0 # Fill missing values with 0 ) %>% select(-category) # Remove the original column print(df_transformed)
This will turn each unique value in category into its own column (with 1s where the row matched the category, 0s otherwise) and drop the original category column.
Using fastDummies (simpler one-liner)
If you prefer a more concise method, the fastDummies package does this in one step, auto-removing the original column:
library(fastDummies) df_transformed <- df %>% dummy_cols(select_columns = "category", remove_selected_columns = TRUE)
gear, split gear unique values into separate columns We'll build on the same logic, but ensure all other columns are preserved while transforming the gear column:
Using tidyverse
library(tidyverse) # Transform mtcars mtcars_transformed <- mtcars %>% rownames_to_column("car_model") %>% # Optional: retain car names as a column (remove if not needed) mutate(indicator = 1) %>% pivot_wider( names_from = gear, values_from = indicator, values_fill = 0, names_prefix = "gear_" # Add a prefix to new column names for clarity ) %>% select(-gear) # Remove the original gear column # Check the first few rows head(mtcars_transformed)
The names_prefix argument ensures new columns are named gear_3, gear_4, gear_5 instead of just numbers, making your data easier to read.
Using fastDummies (quick alternative)
Again, fastDummies simplifies this task drastically:
library(fastDummies) mtcars_transformed <- mtcars %>% dummy_cols(select_columns = "gear", remove_selected_columns = TRUE, prefix = "gear")
This will generate the indicator columns for gear values, drop the original gear column, and keep all other columns from mtcars intact.
内容的提问来源于stack exchange,提问作者Username

