R语言新手求助:使用dplyr处理数据框两列计算问题
Hey there! As someone who’s worked with gene expression data in R quite a bit, I totally get how frustrating it can be when you’re starting out with dplyr and can’t find clear answers online. Let’s walk through common, useful tasks you might want to perform on your immgen_dat tibble—since you mentioned it has sample columns starting with X. and rows representing gene expression values.
First, let’s confirm your data structure to align:
- Non-sample metadata columns:
ProbeSetID,GeneName,Description - Sample expression columns: All columns starting with
X.(likeX.proB_CLP_BM.,X.proB_CLP_FL., etc.)
Here are practical dplyr workflows tailored to this data type:
If you want to keep only genes that have an expression value above a certain cutoff (e.g., 10) in at least one sample:
library(dplyr) high_expression_genes <- immgen_dat %>% rowwise() %>% # Treat each gene row as an individual group mutate(max_expression = max(c_across(starts_with("X.")))) %>% # Calculate max expr per gene filter(max_expression > 10) %>% # Keep only genes meeting the threshold ungroup() # Reset row-wise grouping
Gene expression data is often easier to visualize or analyze in long format (one row per gene-sample pair):
library(tidyr) # Works seamlessly with dplyr long_format_dat <- immgen_dat %>% pivot_longer( cols = starts_with("X."), # Target all sample columns names_to = "Sample", # Name for the new sample column values_to = "Expression" # Name for the new expression value column )
Now your data will have columns: ProbeSetID, GeneName, Description, Sample, Expression
If you want to compute metrics like mean, median, or standard deviation across all samples for each gene:
gene_summary_stats <- immgen_dat %>% rowwise() %>% mutate( mean_expr = mean(c_across(starts_with("X.")), na.rm = TRUE), median_expr = median(c_across(starts_with("X.")), na.rm = TRUE), sd_expr = sd(c_across(starts_with("X.")), na.rm = TRUE) ) %>% ungroup() %>% select(ProbeSetID, GeneName, mean_expr, median_expr, sd_expr) # Keep only relevant columns
The na.rm = TRUE flag ensures missing values don’t break your calculations.
If you need to isolate data for a list of target genes (e.g., "Cd3e" and "Il2ra"):
target_genes <- immgen_dat %>% filter(GeneName %in% c("Cd3e", "Il2ra"))
To standardize expression so each gene has a mean of 0 and standard deviation of 1:
normalized_expression <- immgen_dat %>% rowwise() %>% mutate(across(starts_with("X."), ~ scale(.x)[,1])) %>% # Apply z-score to each sample column ungroup()
If you had a specific task in mind that isn’t covered here—like grouping samples by tissue type or comparing expression between sample groups—just share the details and I can help refine the code!
内容的提问来源于stack exchange,提问作者Atakan

