如何在R语言中基于指定列的值进行数据抽样?
Hey there! Let's walk through how to perform sampling operations based on specific column values using your dataset. First, let's fix the truncated code to get the full reproducible dataset:
df <- data.frame( ID = c("A", "A", "A", "B", "B", "B", "C", "C", "C", "D", "D", "D", "E", "E", "E"), X = c(2L, 4L, 4L, 5L, 3L, 1L, 5L, 2L, 5L, 2L, 3L, 3L, 5L, 4L, 3L), Y = c(2L, 3L, 3L, 3L, 4L, 5L, 3L, 3L, 3L, 3L, 2L, 2L, 3L, 3L, 4L), Z = c(5L, 2L, 2L, 1L, 2L, 3L, 1L, 4L, 1L, 4L, 4L, 4L, 1L, 2L, 2L), Average = c(3L, 3L, 3L, 3L, 3L, 3L, 3L, 3L, 3L, 3L, 3L, 3L, 3L, 3L, 3L) )
Below are common sampling scenarios tailored to your data:
1. Grouped Sampling (e.g., by ID column)
If you want to sample a fixed number or fraction of rows per group (like each ID), the dplyr package makes this straightforward:
Sample a fixed number of rows per group
library(dplyr) # Randomly select 1 row for each ID grouped_sample_n <- df %>% group_by(ID) %>% sample_n(size = 1) %>% # Change size to your desired number ungroup() # View the result print(grouped_sample_n)
Sample a fraction of rows per group
# Select 2/3 of rows for each ID (since each ID has 3 rows, this picks 2 rows) grouped_sample_frac <- df %>% group_by(ID) %>% sample_frac(size = 2/3) %>% # Adjust fraction as needed ungroup()
2. Conditional Sampling (based on column value filters)
If you want to sample rows that meet a specific condition (e.g., X > 3), combine filtering with sampling:
# First filter rows where X > 3, then randomly sample 5 rows from the result conditional_sample <- df %>% filter(X > 3) %>% sample_n(size = 5) # Or if you want all rows matching the condition, just use filter(): all_matching_rows <- df %>% filter(X > 3)
3. Global Sampling (filter first, then sample across the dataset)
Even though all your Average values are 3, here's how you'd sample from rows matching a column value globally:
# Randomly sample 10 rows where Average equals 3 global_sample <- df %>% filter(Average == 3) %>% sample_n(size = 10)
4. Base R Alternative (no packages needed)
If you prefer not to use dplyr, you can achieve grouped sampling with base R functions:
# Randomly select 1 row per ID using split + lapply + rbind base_r_grouped_sample <- do.call( rbind, lapply(split(df, df$ID), function(group) group[sample(nrow(group), 1), ]) )
内容的提问来源于stack exchange,提问作者Geet

