You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从Hmisc包的describe函数提取值并转为tidy格式以用于绘图?

嘿,这个问题我太熟悉了——Hmisc::describe()的输出虽然信息丰富,但确实不是咱们ggplot这类工具喜欢的tidy格式。不过别担心,咱们可以用tidyverse工具把它拆解、重组,变成适合绘图的结构。下面一步步来:

第一步:加载必要的包

首先得把需要的工具包装上(如果没装的话)并加载:

# 安装包(如果没装过)
install.packages(c("Hmisc", "tidyverse"))

# 加载包
library(Hmisc)
library(tidyverse)

第二步:生成describe对象

先把你的数据框(假设叫df)喂给describe(),把结果存成一个对象:

desc_obj <- describe(df)

这个desc_obj其实是个列表,每个元素对应你数据框里的一列,包含了该列的所有统计信息。

第三步:编写提取函数

我们需要一个小函数,把单个变量的describe结果转换成tidy格式。这个函数会自动区分数值型和分类变量,提取对应的统计量:

extract_desc_stats <- function(var_desc) {
  var_name <- var_desc$name
  
  # 处理数值型变量:提取基本统计量和分位数
  if (!is.null(var_desc$mean)) {
    # 基本统计(样本量、缺失值、均值、中位数)
    basic_stats <- tibble(
      variable = var_name,
      n = var_desc$n,
      n_missing = var_desc$nmiss,
      mean = var_desc$mean,
      median = var_desc$median
    )
    # 分位数信息
    quantiles <- as_tibble(t(var_desc$quantiles), rownames = "quantile") %>%
      rename(value = V1) %>%
      mutate(variable = var_name) %>%
      select(variable, quantile, value)
    
    return(list(basic = basic_stats, quantiles = quantiles))
  } 
  # 处理分类变量:提取频率和百分比
  else {
    freq_table <- as_tibble(var_desc$counts, rownames = "category") %>%
      rename(frequency = V1, percent = V2) %>%
      mutate(variable = var_name) %>%
      select(variable, category, frequency, percent)
    
    return(list(frequencies = freq_table))
  }
}

第四步:批量提取并整理成tidy数据框

用purrr::map()批量处理所有变量,然后把结果合并成统一的tidy数据框:

# 对每个变量应用提取函数
desc_list <- map(desc_obj, extract_desc_stats)

# 合并数值变量的基本统计量
basic_tidy <- map_dfr(desc_list, ~.$basic, .default = tibble())

# 合并数值变量的分位数
quantiles_tidy <- map_dfr(desc_list, ~.$quantiles, .default = tibble())

# 合并分类变量的频率表
freq_tidy <- map_dfr(desc_list, ~.$frequencies, .default = tibble())

第五步:用tidy数据绘图

现在这些tidy数据就可以直接喂给ggplot了,举几个常用的例子:

例1:绘制各变量的缺失值数量

ggplot(basic_tidy, aes(x = reorder(variable, n_missing), y = n_missing)) +
  geom_col(fill = "#2E86AB") +
  coord_flip() +
  labs(title = "Missing Values by Variable",
       x = "Variable",
       y = "Number of Missing Values") +
  theme_minimal()

例2:绘制数值变量的分位数分布

ggplot(quantiles_tidy, aes(x = variable, y = value, color = quantile)) +
  geom_point(size = 2) +
  coord_flip() +
  labs(title = "Quantile Distribution by Numeric Variable",
       x = "Variable",
       y = "Value") +
  theme_minimal()

例3:绘制分类变量的频率分布

ggplot(freq_tidy, aes(x = variable, y = frequency, fill = category)) +
  geom_col(position = "dodge", alpha = 0.8) +
  coord_flip() +
  labs(title = "Category Frequencies by Categorical Variable",
       x = "Variable",
       y = "Frequency") +
  theme_minimal()

额外技巧:提取极端值

如果需要提取describe()里的极端值信息,也可以在提取函数的数值变量分支里加上这段:

# 提取极端值
extremes <- as_tibble(var_desc$extremes, rownames = "type") %>%
  rename(value = V1) %>%
  mutate(variable = var_name) %>%
  select(variable, type, value)

然后把它加入返回的列表,后续合并成extremes_tidy即可。


内容的提问来源于stack exchange,提问作者reubenmcg

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 04:06:09