You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于不同分组条件筛选R语言中的DataFrame?

Conditional Filtering by Grouped Ranges in R (dplyr)

Hey there! Let's work through this conditional filtering problem step by step. Since you need different MG selection rules based on specific ranges of loc.id, dplyr's case_when() function is perfect here—it lets you define clear, scenario-based conditions directly inside filter().

Step 1: Set up your data (with reproducibility)

First, let's recreate your data frame with a seed so the random x values are consistent every time you run it:

library(dplyr)

set.seed(123) # Ensures runif() gives the same values each time
df <- data.frame(
  loc.id = rep(1:10, each = 10),
  MG = rep(1:10, times = 10),
  x = runif(100)
)

Step 2: Apply the conditional filter

Here's the code to implement your exact requirements:

filtered_df <- df %>%
  filter(
    case_when(
      # When loc.id is less than 4: keep MG 1-4
      loc.id < 4 ~ MG %in% 1:4,
      # When loc.id is between 5 and 6 (inclusive): keep MG 5-8
      between(loc.id, 5, 6) ~ MG %in% 5:8,
      # When loc.id is greater than 6: keep MG >8 (i.e., 9-10)
      loc.id > 6 ~ MG > 8,
      # Catch-all: exclude any rows that don't match the above (optional here)
      TRUE ~ FALSE
    )
  )

How this works:

  • case_when() evaluates each row against the conditions in order. For each row, it returns TRUE or FALSE based on the first matching condition.
  • MG %in% 1:4 checks if the MG value falls exactly within the 1-4 range (inclusive)—this is cleaner than writing MG >=1 & MG <=4.
  • between(loc.id, 5,6) is a handy dplyr function to check if a value is within a closed range (no need for loc.id >=5 & loc.id <=6).
  • The final TRUE ~ FALSE line is a safety net to exclude any rows that don't fit your defined loc.id ranges (though in your data, loc.id only goes 1-10, so this is optional, but good practice for robustness).

Verify the results

To make sure the filtering worked as expected, you can check the unique combinations of loc.id and MG in the filtered data:

filtered_df %>%
  distinct(loc.id, MG) %>%
  arrange(loc.id, MG)

This will show you exactly which pairs are kept, confirming that:

  • loc.id 1-3 only have MG 1-4
  • loc.id 5-6 only have MG 5-8
  • loc.id7-10 only have MG 9-10

内容的提问来源于stack exchange,提问作者89_Simple

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 08:14:44