You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用R提取数据框各列(含数值与字符串)唯一值并合并为单列?

Extract Unique Values from All Columns into a Two-Column Format

Hey there! Let's figure out how to extract unique values from every column in your data frame (whether they're numeric or character) and combine them into that clean two-column format you want. This solution will work no matter what column names or data types you're dealing with—perfect for any data frame you throw at it.

First, let's fix up your example data a bit: using cbind() creates a matrix instead of a data frame, which can accidentally convert numeric columns to character. Let's use data.frame() instead to keep our data types intact:

a = c("a", "b", "c", "d", "a")
b = c(1, 2, 3, 4, 3)
df <- data.frame(a, b, stringsAsFactors = FALSE)

The tidyverse (specifically dplyr and tidyr) makes this task super clean and readable. Here's how to do it:

library(tidyverse)

result <- df %>%
  # Convert wide data to long format
  pivot_longer(cols = everything(), names_to = "variable", values_to = "Level") %>%
  # Keep only unique combinations of variable and Level
  distinct(variable, Level) %>%
  # Sort to match your desired output
  arrange(variable, Level)

print(result)

What's happening here:

  • pivot_longer(cols = everything()) takes every column in your data frame and converts it into two columns: variable (the original column name) and Level (the values from that column).
  • distinct(variable, Level) removes any duplicate entries so we only keep unique values per column.
  • arrange(variable, Level) sorts the result to match the order you showed in your example.

Method 2: Using Purrr (Another Tidyverse Approach)

If you prefer a more iterative approach, you can use purrr to loop through each column and build your result:

library(purrr)

result <- map_dfr(names(df), function(col_name) {
  # Create a tibble for each column with its name and unique values
  tibble(variable = col_name, Level = unique(df[[col_name]]))
}) %>%
  arrange(variable, Level)

print(result)

What's happening here:

  • map_dfr() loops over each column name, runs the function inside, and binds all the results into a single data frame.
  • For each column, we create a small table with the column name and its unique values, then combine them all together.

Method 3: Base R (No Packages Needed)

If you don't want to use external packages, you can do this with base R functions:

# Create a list of data frames, one for each column's unique values
result_list <- lapply(names(df), function(col_name) {
  data.frame(
    variable = col_name,
    Level = unique(df[[col_name]]),
    stringsAsFactors = FALSE
  )
})

# Combine all the data frames into one
result <- do.call(rbind, result_list)

# Sort and reset row names
result <- result[order(result$variable, result$Level), ]
row.names(result) <- NULL

print(result)

What's happening here:

  • lapply() loops through each column name and creates a data frame for each column's unique values.
  • do.call(rbind, result_list) combines all those small data frames into one big one.
  • Finally, we sort the result and reset the row names to make it clean.

All three methods will give you exactly the output you want:

variable Level
1        a     a
2        a     b
3        a     c
4        a     d
5        b     1
6        b     2
7        b     3
8        b     4

These solutions work for any data frame, regardless of column names or data types (numeric, character, factor—they'll all be handled correctly).

内容的提问来源于stack exchange,提问作者Jordan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 04:12:58