You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言中使用dplyr的filter实现固定范围分层筛选的问题

解决固定范围分层筛选的方法

核心问题原因

你遇到的范围不一致问题,本质是每次筛选时实时计算分位数依赖当前动态子集,如果数据存在重复值或分位数计算的参数差异(比如quantile函数的type参数),会导致阈值波动。要固定范围,必须先基于完整的目标群体(全体男性)预计算好身高、体重的分位数阈值,再用固定值做筛选。

具体实现步骤与代码

1. 预计算固定阈值

先从全量数据中提取男性群体,分别计算身高的20%分位数,以及身高筛选后男性群体的体重20%分位数,将这些阈值存为独立变量:

library(dplyr)

# 假设数据集名为patient_data,包含gender, height, weight, visit_times, prevalence字段
# 提取男性数据集
male_data <- patient_data %>% filter(gender == "男")

# 计算身高20%分位数阈值(固定值)
height_threshold <- quantile(male_data$height, probs = 0.2, na.rm = TRUE)

# 筛选身高最矮20%的男性,再计算体重20%分位数阈值(固定值)
weight_threshold <- male_data %>%
  filter(height <= height_threshold) %>%
  pull(weight) %>%
  quantile(probs = 0.2, na.rm = TRUE)

2. 用固定阈值筛选目标人群

基于预计算的阈值,筛选出最终目标群体:

target_group <- male_data %>%
  filter(height <= height_threshold,
         weight <= weight_threshold)

3. 按就医次数分位数区间统计患病率

将就医次数划分为5个等距分位数区间,然后统计各区间的患病率:

result <- target_group %>%
  # 生成就医次数的分位数区间标签
  mutate(visit_interval = cut(visit_times,
                             breaks = quantile(visit_times, probs = seq(0, 1, 0.2), na.rm = TRUE),
                             include.lowest = TRUE,
                             labels = c("0-20%", "20%-40%", "40%-60%", "60%-80%", "80%-100%"))) %>%
  # 按区间分组统计患病率
  group_by(visit_interval) %>%
  summarise(
    total_count = n(),
    prevalence_rate = mean(prevalence, na.rm = TRUE) * 100  # 假设prevalence是0/1变量,转为百分比
  ) %>%
  ungroup()

关键注意点

  • 计算分位数时,务必添加na.rm = TRUE处理缺失值,避免因缺失值导致阈值计算错误
  • 如果分位数区间出现重叠或空组,可调整cut函数的breaks参数,比如用unique(quantile(...))去重,或手动指定区间边界
  • 若需要严格的分位数分组(保证每组数量尽可能均等),可以用ntile函数替代cut:
    mutate(visit_ntile = ntile(visit_times, n = 5)) %>%
    group_by(visit_ntile)
    

内容的提问来源于stack exchange,提问作者stats_noob

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.17 14:42:43