You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言按分组替换数据框NA:数值列用均值、字符列用众数

问题根因
  • num.of.doors列填充值错误:R基础环境没有内置计算众数的mode()函数,你调用的mode()是查询对象存储类型的基础函数,传入字符向量时会直接返回"character",并非列的实际众数,因此填充结果不符合预期。
  • for循环运行无效:循环中的i是从1到总列数的整数索引,判断is.numeric(i)的结果永远为真,完全没有读取列的实际内容;同时replace_na()不会原地修改输入对象,你没有将替换后的结果赋值回原数据集对应列,因此运行后数据不会发生任何变化。
非分组基础填充实现

首先自定义众数计算函数(R无内置众数函数):

# 计算向量众数,自动忽略NA值
get_mode <- function(x) {
  x_non_na <- na.omit(x)
  unique_vals <- unique(x_non_na)
  unique_vals[which.max(tabulate(match(x_non_na, unique_vals)))]
}

再编写正确的批量填充逻辑,需要先加载tidyr包以使用replace_na:

library(tidyr)
for (i in seq_along(auto_data)) {
  current_col <- auto_data[[i]]
  # 数值列用均值填充NA
  if (is.numeric(current_col)) {
    auto_data[[i]] <- replace_na(current_col, mean(current_col, na.rm = TRUE))
  }
  # 字符/因子列用众数填充NA
  if (is.character(current_col) | is.factor(current_col)) {
    auto_data[[i]] <- replace_na(current_col, get_mode(current_col))
  }
}

逻辑说明:用[[按索引提取数据集列的实际内容,判断类型后完成填充,再将结果赋值回原列,即可直接修改原数据集。

按make、body_style分组填充实现

加载dplyr包用分组语法实现,无需手动处理分组切片,代码更简洁:

library(dplyr)
library(tidyr)

auto_data_filled <- auto_data %>%
  group_by(make, body.style) %>%
  mutate(
    # 所有数值列使用分组均值填充NA
    across(where(is.numeric), ~replace_na(.x, mean(.x, na.rm = TRUE))),
    # 所有字符/因子列使用分组众数填充NA
    across(where(~is.character(.x) | is.factor(.x)), ~replace_na(.x, get_mode(.x)))
  ) %>%
  ungroup()

如果存在分组内某列全为NA、无法计算分组均值/众数的边界情况,可以增加全局统计量兜底逻辑,避免出现NaN或NULL填充:

# 带兜底逻辑的众数计算
get_mode <- function(x) {
  x_non_na <- na.omit(x)
  if (length(x_non_na) == 0) return(NA)
  unique_vals <- unique(x_non_na)
  unique_vals[which.max(tabulate(match(x_non_na, unique_vals)))]
}

auto_data_filled <- auto_data %>%
  # 提前计算全局均值、全局众数作为兜底值
  mutate(
    across(where(is.numeric), ~mean(.x, na.rm = TRUE), .names = "global_mean_{.col}"),
    across(where(~is.character(.x) | is.factor(.x)), ~get_mode(.x), .names = "global_mode_{.col}")
  ) %>%
  group_by(make, body.style) %>%
  mutate(
    across(where(is.numeric), function(col) {
      grp_mean <- mean(col, na.rm = TRUE)
      fill_val <- ifelse(is.nan(grp_mean), cur_data()[[paste0("global_mean_", cur_column())]], grp_mean)
      ifelse(is.na(col), fill_val, col)
    }),
    across(where(~is.character(.x) | is.factor(.x)), function(col) {
      grp_mode <- get_mode(col)
      fill_val <- ifelse(is.na(grp_mode), cur_data()[[paste0("global_mode_", cur_column())]], grp_mode)
      ifelse(is.na(col), fill_val, col)
    })
  ) %>%
  ungroup() %>%
  # 删除临时生成的全局统计量列
  select(-starts_with("global_mean_"), -starts_with("global_mode_"))

内容的提问来源于stack exchange,提问作者Abbi Asseged

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.27 23:27:10