You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

按类别均值插补缺失值遇问题:代码仅返回缺失值未计算均值

按类别自动插补缺失值的解决方案

你用aggregate得到全缺失值,大概率是这几个原因:

  • 某类别下对应变量的所有值都是NA,均值自然为NA
  • 待计算的变量不是数值型(比如因子/字符型),mean函数无法返回有效结果
  • 分组变量Company_category本身包含NA,这些NA会被单独作为无效分组,对应结果全是NA

下面是两种自动完成按类别插补的可行方案:

方案一:用dplyr(代码更简洁)

library(dplyr)

# 1. 按类别计算所有数值变量的均值(忽略NA,排除NA分组)
mean_per_class <- df %>%
  filter(!is.na(Company_category)) %>%
  group_by(Company_category) %>%
  summarise(across(where(is.numeric), ~mean(.x, na.rm = TRUE)), .groups = "drop")

# 2. 自动对原数据的缺失值进行插补
df_imputed <- df %>%
  left_join(mean_per_class, by = "Company_category") %>%
  # 遍历数值变量,将NA替换为对应类别的均值
  mutate(across(ends_with(".x"), 
                ~ifelse(is.na(.x), get(sub(".x", ".y", cur_column())), .x)),
         .keep = "unused") %>%
  # 还原变量名
  rename_with(~sub(".x", "", .x), ends_with(".x"))

方案二:用Base R(无需额外安装包)

# 1. 按类别计算数值变量的均值,跳过非数值变量,排除NA分组
mean_per_class <- aggregate(. ~ Company_category, data = df,
                            FUN = function(x) {
                              if (is.numeric(x)) mean(x, na.rm = TRUE) else x
                            },
                            na.action = na.omit)

# 2. 自动遍历所有数值变量完成插补
df_imputed <- df
# 筛选出需要处理的数值变量(跳过分组变量)
target_cols <- setdiff(names(df_imputed), "Company_category")
target_cols <- target_cols[sapply(df_imputed[target_cols], is.numeric)]

for (col in target_cols) {
  # 匹配观测所属类别,替换NA为对应均值
  df_imputed[[col]] <- ifelse(
    is.na(df_imputed[[col]]),
    mean_per_class[match(df_imputed$Company_category, mean_per_class$Company_category), col],
    df_imputed[[col]]
  )
}

关键说明

  • 只对数值型变量做均值插补,非数值变量(如字符、因子)自动跳过;如果需要插补这类变量,可以把mean换成众数计算逻辑(比如function(x) names(which.max(table(x))))
  • 代码自动遍历所有符合条件的变量,无需手动逐个指定
  • 已处理分组变量含NA的情况,避免无效分组导致的全NA结果

内容的提问来源于stack exchange,提问作者Simone Carminati

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.27 09:52:39