You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

ggplot如何固定指定变量颜色、其余随机,并自动展示Top5高频项

解决方案

数据预处理:自动保留目标类别

首先对数据进行处理,自动筛选出text列中的"a",以及除"a"外出现频率最高的前5个类别,其余类别可统一归为"Other"(若不需要合并其余类别,也可直接过滤):

library(tidyverse)

# 原始数据(用户提供的代码)
df <- setNames(
  data.frame(
    as.POSIXct(
      c(
        "2022-07-29 00:00:00", "2022-07-29 00:00:05",
        "2022-07-29 00:05:00", "2022-07-29 00:05:05",
        "2022-07-29 00:10:00", "2022-07-29 00:15:00",
        "2022-07-29 00:20:00", "2022-07-29 00:20:05"
      )
    ),
    c(1, 2, 3, 4, 5, 6, 7, 8),
    c("a", "a", "b847", "b317", "b317", "bob680", "bf456", "c3400")
  ),
  c("timeStamp", "value1", "text")
)

# 数据预处理:保留"a"和除它外频率最高的前5个类别,其余归为"Other"
df_processed <- df %>%
  # 计算每个text的出现频率
  count(text, name = "freq") %>%
  # 按频率降序排列
  arrange(desc(freq)) %>%
  # 提取"a"和其余top5的text值
  slice(
    c(which(text == "a"), # 定位"a"的位置
       (1:n())[text != "a"][1:5]) # 取非"a"的前5个
  ) %>%
  pull(text) %>%
  # 将原始数据中不在保留列表的text转为"Other"
  { keep_text <- .; df %>% mutate(text = ifelse(text %in% keep_text, text, "Other")) }

自定义配色:固定"a"为红色,其余自动配色

使用scale_fill_manual固定"a"为红色,其余类别自动使用ggplot的默认色调,无需手动指定所有未知类别:

# 生成可视化图表
df_processed %>%
  ggplot(aes(x = fct_infreq(text), fill = text)) +
  geom_bar(aes(y = (..count..)/sum(..count..)), stat = "count") +
  # 自定义填充颜色:"a"设为红色,其余用默认色调
  scale_fill_manual(
    values = c(
      "a" = "red",
      setNames(
        scales::hue_pal()(length(setdiff(unique(df_processed$text), "a"))),
        setdiff(unique(df_processed$text), "a")
      )
    )
  ) +
  labs(x = "Text类别", y = "占比", title = "文本类别频率占比") # 可选:添加标签

关键说明

  • 数据预处理部分通过count()统计频率,slice()精准筛选目标类别,实现完全自动化,无需手动识别top5;
  • 配色部分利用scales::hue_pal()生成ggplot默认的色调,仅替换"a"的颜色,适配任意未知的其余类别;
  • 若不需要合并其余类别为"Other",可将预处理最后一步改为df %>% filter(text %in% keep_text),直接过滤掉低频类别。

内容的提问来源于stack exchange,提问作者ambergris

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.13 23:15:47