在R中用简单规则筛选数据框行以优化Class列总和
基于单特征阈值筛选最大化Class总和的解决方案
你的需求本质是通过单特征的阈值规则(如x < threshold)筛选样本,使得保留样本的Class总和最大化——总和越大,说明留下的"好样本"(Class=1)越多、剔除的"坏样本"(Class=-1)越多,同时尽量少剔除好样本。
实现思路
对于每个数值型特征,我们可以:
- 提取该特征的所有唯一值作为候选阈值(或取分位数减少计算量)
- 对每个阈值,计算两种筛选规则(
特征 < 阈值和特征 >= 阈值)下的Class总和 - 记录每个特征对应的最优阈值、最优规则及对应的最大总和
- 最终选择所有特征中总和最大的那个规则作为筛选条件
代码实现
# 加载数据(你的代码) pacman::p_load("leaps", "tidyverse","caret","magrittr") data("GermanCredit") GermanCredit <- GermanCredit %>% mutate(Class=if_else(Class=="Good",1,-1)) # 定义函数:计算单个特征的最优阈值及最大总和 find_best_threshold <- function(data, feature_col, target_col = "Class") { feature_vals <- data[[feature_col]] # 只处理数值型特征 if (!is.numeric(feature_vals)) { return(tibble(feature = feature_col, threshold = NA, rule = NA, max_sum = NA)) } # 提取候选阈值(去重并排序) thresholds <- sort(unique(feature_vals)) # 如果特征值太少,直接返回 if (length(thresholds) < 2) { return(tibble(feature = feature_col, threshold = NA, rule = NA, max_sum = NA)) } # 遍历所有阈值,计算两种规则的总和 results <- map_dfr(thresholds, function(thresh) { sum_lt <- data %>% filter(.data[[feature_col]] < thresh) %>% pull(target_col) %>% sum() sum_ge <- data %>% filter(.data[[feature_col]] >= thresh) %>% pull(target_col) %>% sum() tibble( threshold = thresh, rule = ifelse(sum_lt > sum_ge, paste0(feature_col, " < ", thresh), paste0(feature_col, " >= ", thresh)), total_sum = max(sum_lt, sum_ge) ) }) # 找到该特征下的最优结果 best_result <- results %>% slice_max(total_sum, n = 1) %>% mutate(feature = feature_col) select(best_result, feature, threshold, rule, total_sum) } # 对所有特征计算最优阈值 all_features <- names(GermanCredit) best_thresholds <- map_dfr(all_features, ~find_best_threshold(GermanCredit, .x)) # 过滤掉无效结果(非数值特征),并按总和降序排列 best_thresholds_valid <- best_thresholds %>% filter(!is.na(total_sum)) %>% arrange(desc(total_sum)) # 查看最优的前5个规则 print(best_thresholds_valid %>% head(5)) # 应用最优规则筛选数据 best_rule <- best_thresholds_valid$rule[1] filtered_data <- GermanCredit %>% filter(eval(parse(text = best_rule))) # 查看筛选前后的总和对比 cat("原始总和:", sum(GermanCredit$Class), "\n") cat("筛选后总和:", sum(filtered_data$Class), "\n") cat("筛选规则:", best_rule, "\n")
结果解释
- 运行代码后,
best_thresholds_valid会展示所有数值特征的最优规则及对应的总和,排在最前面的就是效果最好的单特征筛选规则 - 比如可能得到类似
Duration < 12这样的规则,意味着保留贷款时长小于12个月的样本,能让Class总和最大化 - 你可以根据实际需求调整规则(比如如果某个规则剔除的样本太多,即使总和高也可以选择次优的)
针对分类特征的补充
如果想对分类特征使用类似规则(比如x == "某类别"),可以修改函数,遍历分类特征的每个类别,计算保留该类别样本的Class总和,选择总和最大的类别作为筛选条件。
内容的提问来源于stack exchange,提问作者gilberke
相关产品推荐
相关产品推荐

