在R中计算基因型产量基础统计量:多地点子集数据统计需求
Hey there! I totally get how frustrating it can be when you're stuck trying to automate these kinds of repeated stats calculations—let's break this down step by step with R code you can plug right into your dataset.
First, let's make sure we're on the same page: I'll assume your data is in a long format (columns like genotype, location, yield). If your data is wide (one column per location), we can reshape it first, but I'll start with the more common long format since it's easier to work with for grouping.
Step 1: Simulate a Sample Dataset (Skip if you have your own data)
Let's create a fake dataset that matches your setup (300 genotypes, 6 locations each) to test the code:
set.seed(123) # Makes results reproducible genotypes <- paste0("G", 1:300) locations <- paste0("Loc", 1:6) # Create long-format data frame df <- expand.grid(genotype = genotypes, location = locations) df$yield <- rnorm(nrow(df), mean = 50, sd = 8) # Simulate yield values
Step 2: Calculate Stats for All 6 Locations
We'll use the dplyr package (part of the tidyverse) because it makes grouping and summarizing super intuitive. If you don't have it installed, run install.packages("dplyr") first.
library(dplyr) # Stats for all 6 locations per genotype stats_6loc <- df %>% group_by(genotype) %>% summarise( num_locations = n(), mean_yield = mean(yield, na.rm = TRUE), sd_yield = sd(yield, na.rm = TRUE), se_yield = sd_yield / sqrt(num_locations) # Standard error = SD / sqrt(n) ) # Check the first few rows head(stats_6loc)
Step 3: Calculate Stats for 5 Locations per Genotype
Here, I'll assume you want to randomly select 5 locations per genotype (if you need to exclude specific locations instead, just swap out the sample_n() line with a filter(location != "LocX") where LocX is the location to remove).
# Stats for random 5 locations per genotype stats_5loc <- df %>% group_by(genotype) %>% sample_n(size = 5) %>% # Pick 5 random locations for each genotype summarise( num_locations = n(), mean_yield = mean(yield, na.rm = TRUE), sd_yield = sd(yield, na.rm = TRUE), se_yield = sd_yield / sqrt(num_locations) ) head(stats_5loc)
Step 4: Calculate Stats for 4 Locations per Genotype
Same logic as above, just sample 4 locations instead:
# Stats for random 4 locations per genotype stats_4loc <- df %>% group_by(genotype) %>% sample_n(size = 4) %>% summarise( num_locations = n(), mean_yield = mean(yield, na.rm = TRUE), sd_yield = sd(yield, na.rm = TRUE), se_yield = sd_yield / sqrt(num_locations) ) head(stats_4loc)
Bonus: If You Need All Possible Location Combinations
If you don't want random samples—instead you want stats for every possible set of 5 or 4 locations (e.g., all 6 possible 5-location combinations per genotype)—use the purrr package to handle combinations:
library(purrr) # Get stats for every 5-location combination per genotype all_5loc_combo_stats <- df %>% group_by(genotype) %>% group_map(function(.x, .y) { # Get all locations for the current genotype all_locs <- unique(.x$location) # Generate all possible 5-location combinations loc_combos <- combn(all_locs, 5, simplify = FALSE) # Calculate stats for each combination map_dfr(loc_combos, function(combo) { subset_data <- .x %>% filter(location %in% combo) tibble( genotype = .y$genotype, selected_locations = paste(combo, collapse = ", "), num_locations = nrow(subset_data), mean_yield = mean(subset_data$yield, na.rm = TRUE), sd_yield = sd(subset_data$yield, na.rm = TRUE), se_yield = sd_yield / sqrt(num_locations) ) }) }) %>% bind_rows() # Check results for genotype G1 filter(all_5loc_combo_stats, genotype == "G1")
Combine All Results Into One Table
To make it easy to compare stats across 6, 5, and 4 locations, bind the results together with a label column:
combined_stats <- bind_rows( stats_6loc %>% mutate(location_set = "6 locations"), stats_5loc %>% mutate(location_set = "5 locations"), stats_4loc %>% mutate(location_set = "4 locations") ) %>% select(location_set, genotype, everything()) head(combined_stats)
Quick Notes:
- If your data has missing values, the
na.rm = TRUEinmean()andsd()ensures those are ignored. - If your data is wide (one column per location), use
tidyr::pivot_longer()to reshape it to long format first. - If you prefer base R over tidyverse, you can use
aggregate()or loops, but tidyverse is way more readable for this kind of repeated grouping.
内容的提问来源于stack exchange,提问作者Tyler Durden

