如何用R提取数据框各列(含数值与字符串)唯一值并合并为单列?
Hey there! Let's figure out how to extract unique values from every column in your data frame (whether they're numeric or character) and combine them into that clean two-column format you want. This solution will work no matter what column names or data types you're dealing with—perfect for any data frame you throw at it.
First, let's fix up your example data a bit: using cbind() creates a matrix instead of a data frame, which can accidentally convert numeric columns to character. Let's use data.frame() instead to keep our data types intact:
a = c("a", "b", "c", "d", "a") b = c(1, 2, 3, 4, 3) df <- data.frame(a, b, stringsAsFactors = FALSE)
Method 1: Using Tidyverse (Recommended)
The tidyverse (specifically dplyr and tidyr) makes this task super clean and readable. Here's how to do it:
library(tidyverse) result <- df %>% # Convert wide data to long format pivot_longer(cols = everything(), names_to = "variable", values_to = "Level") %>% # Keep only unique combinations of variable and Level distinct(variable, Level) %>% # Sort to match your desired output arrange(variable, Level) print(result)
What's happening here:
pivot_longer(cols = everything())takes every column in your data frame and converts it into two columns:variable(the original column name) andLevel(the values from that column).distinct(variable, Level)removes any duplicate entries so we only keep unique values per column.arrange(variable, Level)sorts the result to match the order you showed in your example.
Method 2: Using Purrr (Another Tidyverse Approach)
If you prefer a more iterative approach, you can use purrr to loop through each column and build your result:
library(purrr) result <- map_dfr(names(df), function(col_name) { # Create a tibble for each column with its name and unique values tibble(variable = col_name, Level = unique(df[[col_name]])) }) %>% arrange(variable, Level) print(result)
What's happening here:
map_dfr()loops over each column name, runs the function inside, and binds all the results into a single data frame.- For each column, we create a small table with the column name and its unique values, then combine them all together.
Method 3: Base R (No Packages Needed)
If you don't want to use external packages, you can do this with base R functions:
# Create a list of data frames, one for each column's unique values result_list <- lapply(names(df), function(col_name) { data.frame( variable = col_name, Level = unique(df[[col_name]]), stringsAsFactors = FALSE ) }) # Combine all the data frames into one result <- do.call(rbind, result_list) # Sort and reset row names result <- result[order(result$variable, result$Level), ] row.names(result) <- NULL print(result)
What's happening here:
lapply()loops through each column name and creates a data frame for each column's unique values.do.call(rbind, result_list)combines all those small data frames into one big one.- Finally, we sort the result and reset the row names to make it clean.
All three methods will give you exactly the output you want:
variable Level 1 a a 2 a b 3 a c 4 a d 5 b 1 6 b 2 7 b 3 8 b 4
These solutions work for any data frame, regardless of column names or data types (numeric, character, factor—they'll all be handled correctly).
内容的提问来源于stack exchange,提问作者Jordan

