ggplot如何固定指定变量颜色、其余随机,并自动展示Top5高频项
解决方案
数据预处理:自动保留目标类别
首先对数据进行处理,自动筛选出text列中的"a",以及除"a"外出现频率最高的前5个类别,其余类别可统一归为"Other"(若不需要合并其余类别,也可直接过滤):
library(tidyverse) # 原始数据(用户提供的代码) df <- setNames( data.frame( as.POSIXct( c( "2022-07-29 00:00:00", "2022-07-29 00:00:05", "2022-07-29 00:05:00", "2022-07-29 00:05:05", "2022-07-29 00:10:00", "2022-07-29 00:15:00", "2022-07-29 00:20:00", "2022-07-29 00:20:05" ) ), c(1, 2, 3, 4, 5, 6, 7, 8), c("a", "a", "b847", "b317", "b317", "bob680", "bf456", "c3400") ), c("timeStamp", "value1", "text") ) # 数据预处理:保留"a"和除它外频率最高的前5个类别,其余归为"Other" df_processed <- df %>% # 计算每个text的出现频率 count(text, name = "freq") %>% # 按频率降序排列 arrange(desc(freq)) %>% # 提取"a"和其余top5的text值 slice( c(which(text == "a"), # 定位"a"的位置 (1:n())[text != "a"][1:5]) # 取非"a"的前5个 ) %>% pull(text) %>% # 将原始数据中不在保留列表的text转为"Other" { keep_text <- .; df %>% mutate(text = ifelse(text %in% keep_text, text, "Other")) }
自定义配色:固定"a"为红色,其余自动配色
使用scale_fill_manual固定"a"为红色,其余类别自动使用ggplot的默认色调,无需手动指定所有未知类别:
# 生成可视化图表 df_processed %>% ggplot(aes(x = fct_infreq(text), fill = text)) + geom_bar(aes(y = (..count..)/sum(..count..)), stat = "count") + # 自定义填充颜色:"a"设为红色,其余用默认色调 scale_fill_manual( values = c( "a" = "red", setNames( scales::hue_pal()(length(setdiff(unique(df_processed$text), "a"))), setdiff(unique(df_processed$text), "a") ) ) ) + labs(x = "Text类别", y = "占比", title = "文本类别频率占比") # 可选:添加标签
关键说明
- 数据预处理部分通过
count()统计频率,slice()精准筛选目标类别,实现完全自动化,无需手动识别top5; - 配色部分利用
scales::hue_pal()生成ggplot默认的色调,仅替换"a"的颜色,适配任意未知的其余类别; - 若不需要合并其余类别为"Other",可将预处理最后一步改为
df %>% filter(text %in% keep_text),直接过滤掉低频类别。
内容的提问来源于stack exchange,提问作者ambergris
相关产品推荐
相关产品推荐

