You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何计算变量的聚合均值(mean)与标准差(SD)并可视化?

变量重要性聚合统计与可视化解决方案

一、数据处理:生成包含均值和标准差的汇总表

使用tidyverse工具包完成数据整理与统计,步骤如下:

  1. 加载依赖包并拆分行名
library(tidyverse)

# 将行名转为列,并拆分为Fold和变量名称
importance_df2 <- importance_df2 %>%
  rownames_to_column("id") %>%
  separate(id, into = c("Fold", "Variable"), sep = "\\.")
  1. 转换为长格式并计算聚合统计量
# 转换为长格式,便于分组计算
long_df <- importance_df2 %>%
  pivot_longer(cols = c(setosa, versicolor, virginica), 
               names_to = "Species", 
               values_to = "Importance")

# 按变量和物种分组,计算均值与标准差
summary_df <- long_df %>%
  group_by(Variable, Species) %>%
  summarise(Mean_Importance = mean(Importance),
            SD_Importance = sd(Importance),
            .groups = "drop")

# 整理为4行的宽格式(每个变量一行)
wide_summary <- summary_df %>%
  pivot_wider(names_from = Species,
              values_from = c(Mean_Importance, SD_Importance),
              names_sep = "_")

运行后wide_summary即为你需要的4行数据,每行对应一个变量,包含各物种的重要性均值和标准差。

二、绘制均值误差棒图

使用ggplot2绘制每个物种的变量重要性均值及误差棒(标准差):

ggplot(summary_df, aes(x = Variable, y = Mean_Importance, fill = Species)) +
  # 绘制柱状图展示均值
  geom_col(position = position_dodge(width = 0.8), width = 0.7) +
  # 添加误差棒(上下各1个标准差)
  geom_errorbar(aes(ymin = Mean_Importance - SD_Importance, 
                    ymax = Mean_Importance + SD_Importance),
                position = position_dodge(width = 0.8),
                width = 0.2) +
  # 设置图表标签
  labs(title = "变量重要性均值及标准差",
       x = "变量",
       y = "重要性均值",
       fill = "物种") +
  # 使用简洁主题并调整x轴标签角度
  theme_minimal() +
  theme(axis.text.x = element_text(angle = 45, hjust = 1))

该图表会展示每个变量在三个物种中的重要性均值,误差棒表示1倍标准差,直观呈现数据的离散程度。

内容的提问来源于stack exchange,提问作者SK77

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.22 03:37:32