基于阈值筛选Dataframe列并生成含唯一值的新Dataframe
解决方案:筛选低唯一值列并生成新DataFrame
需求说明
从原DataFrame中筛选出唯一值数量少于5个的列,将这些列的名称和对应唯一值列表整理成新DataFrame,每列对应新DataFrame的一行。
原数据加载
首先加载提供的原数据:
df <- structure(list(id = c(1, 2, 3, 4, 5, 6, 7, 8, 9, 10), gender = c("male", "female", "female", "female", "female", "male", "male", "female", "female", "female"), ranking = c("low", "medium", "medium", "medium", "high", "low", "medium", "low", "low", "low"), comments = c("I was really dissapointed by the fact that there was no response", "I got feedback from them but I considered it a lie", "The feedback was really good and I felt convinced", "I was informed they will get back to me", "The feedback was appropriate to me", "I feel the contact person wasn't knowledgeable about the product", "I was told they will follow up within a week but they failed to", "I liked their customer service", "I was told that the issue will soon be addressed", "I am satisfied with the resonse they gave")), class = c("tbl_df", "tbl", "data.frame"), row.names = c(NA, -10L))
实现方法
方法1:基础R实现
无需额外依赖包,通过遍历列完成筛选和整理:
# 设置阈值 threshold <- 5 # 初始化结果DataFrame result_df <- data.frame( variable = character(), unique_values = character(), stringsAsFactors = FALSE ) # 遍历每一列 for(col_name in names(df)) { col_unique <- unique(df[[col_name]]) # 检查唯一值数量是否小于阈值 if(length(col_unique) < threshold) { # 将唯一值转为逗号分隔的字符串 vals_str <- paste(col_unique, collapse = ", ") # 添加到结果DataFrame result_df <- rbind(result_df, data.frame(variable = col_name, unique_values = vals_str)) } } # 查看结果 print(result_df)
方法2:tidyverse简洁实现
使用dplyr和tibble包,代码更简洁高效:
library(dplyr) library(tibble) threshold <- 5 result_df <- df %>% # 提取每列的唯一值并存储为列表 summarise(across(everything(), ~ list(unique(.x)))) %>% # 宽表转长表,拆分列名和唯一值列表 pivot_longer(everything(), names_to = "variable", values_to = "unique_values") %>% # 计算每列的唯一值数量 mutate(unique_count = lengths(unique_values)) %>% # 筛选唯一值数量小于阈值的行 filter(unique_count < threshold) %>% # 将唯一值列表转为逗号分隔的字符串 mutate(unique_values = sapply(unique_values, paste, collapse = ", ")) %>% # 保留需要的列 select(variable, unique_values) print(result_df)
最终结果
两种方法都会得到如下新DataFrame:
variable unique_values 1 gender male, female 2 ranking low, medium, high
内容的提问来源于stack exchange,提问作者Stephen Okiya
相关产品推荐
相关产品推荐

