You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用R在每行缺失数据占比低于10%时按行均值插补缺失数据

R语言按行均值插补缺失值(仅处理缺失占比<10%的行)

步骤1:准备示例数据

先构造带缺失值的数据集用于演示:

set.seed(123) # 设随机种子保证结果可复现
df <- data.frame(
  col1 = sample(c(1:10, NA), 20, replace = TRUE),
  col2 = sample(c(1:10, NA), 20, replace = TRUE),
  col3 = sample(c(1:10, NA), 20, replace = TRUE),
  col4 = sample(c(1:10, NA), 20, replace = TRUE),
  col5 = sample(c(1:10, NA), 20, replace = TRUE),
  col6 = sample(c(1:10, NA), 20, replace = TRUE),
  col7 = sample(c(1:10, NA), 20, replace = TRUE),
  col8 = sample(c(1:10, NA), 20, replace = TRUE),
  col9 = sample(c(1:10, NA), 20, replace = TRUE),
  col10 = sample(c(1:10, NA), 20, replace = TRUE)
)

步骤2:标记需要处理的行

计算每行缺失值占比,筛选出缺失占比低于10%的行:

# 计算每行缺失占比
row_miss_rate <- rowMeans(is.na(df))
# 生成逻辑向量:TRUE表示该行需要插补
need_impute <- row_miss_rate < 0.1

步骤3:执行均值插补

提供两种主流实现方式:

方式1:Base R 原生实现

复制原数据避免修改原始数据,用apply遍历目标行完成插补:

df_imputed <- df

# 对需要插补的行进行处理
df_imputed[need_impute, ] <- t(apply(df[need_impute, ], 1, function(x) {
  row_mean <- mean(x, na.rm = TRUE) # 计算该行非NA值的均值
  x[is.na(x)] <- row_mean # 替换NA为均值
  return(x)
}))

方式2:tidyverse 工具链实现

适合习惯dplyr语法的用户,用rowwise按行处理:

library(dplyr)

df_imputed_tidy <- df %>%
  rowwise() %>%
  mutate(
    miss_rate = mean(is.na(c_across(everything()))), # 计算当前行缺失占比
    row_mean = mean(c_across(everything()), na.rm = TRUE) # 计算当前行非NA均值
  ) %>%
  # 仅对符合条件的行替换NA
  mutate(across(everything(), ~ ifelse(miss_rate < 0.1 & is.na(.), row_mean, .))) %>%
  select(-miss_rate, -row_mean) %>% # 移除辅助计算列
  ungroup()

验证插补结果

可以对比原数据和插补后的数据,确认效果:

# 查看原数据中需要处理的行
df[need_impute, ]
# 查看插补后的对应行
df_imputed[need_impute, ]

内容的提问来源于stack exchange,提问作者user13751413

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.22 20:54:30