You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

移除NA行后重新调整权重以保持数据集代表性

解决方案:移除NA行并调整权重以保持数据集代表性

要实现移除Class列含NA的行,同时让子集仍能代表原始数据集,核心思路是基于Age和Country分组,让每个组在子集中的总权重与原始数据保持一致——通过放大剩余样本的权重,填补因移除NA行损失的组权重。

示例数据

首先定义你提供的示例数据框:

df <- data.frame(
  Age = c(10, 20, 30, 25, 50, 60, 40),
  Country = c("Germany", "Germany", "Germany", "China", "China", "China", "China"),
  Class = c("A", "B", NA, NA, "B", "A", "A"),
  Weight = c(1.1, 0.8, 1.2, 1.7, 0.7, 1.3, 0.9)
)

方法1:用dplyr(tidyverse)实现

library(dplyr)

# 第一步:计算原始数据中每个(Age, Country)组的总权重
original_group_weights <- df %>%
  group_by(Age, Country) %>%
  summarise(original_total = sum(Weight), .groups = "drop")

# 第二步:过滤NA行并调整权重
adjusted_df <- df %>%
  filter(!is.na(Class)) %>%
  group_by(Age, Country) %>%
  mutate(
    # 计算当前组的总权重
    current_total = sum(Weight),
    # 匹配对应组的原始总权重,计算调整因子
    adjust_factor = original_group_weights$original_total[match(paste(Age, Country), paste(original_group_weights$Age, original_group_weights$Country))] / current_total,
    # 生成调整后的权重
    Adjusted_Weight = Weight * adjust_factor
  ) %>%
  ungroup() %>%
  # 移除中间计算列(可选)
  select(-current_total, -adjust_factor)

方法2:用Base R实现

如果不想依赖tidyverse包,也可以用基础R代码完成:

# 计算原始数据中每个(Age, Country)组的总权重
original_totals <- aggregate(Weight ~ Age + Country, data = df, sum)
names(original_totals)[3] <- "original_total"

# 过滤掉Class为NA的行
filtered_df <- df[!is.na(df$Class), ]

# 计算过滤后每个组的总权重
current_totals <- aggregate(Weight ~ Age + Country, data = filtered_df, sum)
names(current_totals)[3] <- "current_total"

# 计算权重调整因子
adjustment_factors <- merge(original_totals, current_totals, by = c("Age", "Country"))
adjustment_factors$factor <- adjustment_factors$original_total / adjustment_factors$current_total

# 合并因子并计算调整后权重
adjusted_df_base <- merge(filtered_df, adjustment_factors, by = c("Age", "Country"))
adjusted_df_base$Adjusted_Weight <- adjusted_df_base$Weight * adjusted_df_base$factor

# 整理最终列顺序
adjusted_df_base <- adjusted_df_base[, c("Age", "Country", "Class", "Weight", "Adjusted_Weight")]

验证逻辑

以Germany组为例:

  • 原始组总权重:1.1 + 0.8 + 1.2 = 3.1
  • 过滤NA行后剩余样本总权重:1.1 + 0.8 = 1.9
  • 调整因子:3.1 / 1.9 ≈ 1.6316
  • 调整后样本权重:1.1 * 1.6316 ≈ 1.7948,0.8 * 1.6316 ≈ 1.3052,两者总和仍为3.1,与原始组总权重一致。

这样处理后,子集的Age+Country组权重分布完全匹配原始数据,确保了子集的代表性。

内容的提问来源于stack exchange,提问作者Kamaloka

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.31 01:57:37