FireBrowse TCGA LUAD数据不同分期肿瘤样本DEG筛选遇阻求助
Hey there, let's walk through the most likely reasons you're not getting any differentially expressed genes (DEGs) when comparing LUAD tumor stages using FireBrowse's RNASeqV2 RSEM data, and how to fix each issue:
1. Double-Check Your Sample Grouping
First, make sure you're actually selecting the right samples for each stage—FireBrowse's pathologic_stage field uses full labels like "Stage IA", "Stage IB", "Stage IVA", etc., so partial matches might miss some samples. Also, ensure you're not including normal adjacent tissue samples by accident.
Example R Code to Validate Groups:
# Assuming your metadata is stored in a data frame called luad_meta # Filter for strictly tumor samples (exclude normal) tumor_meta <- luad_meta[luad_meta$sample_type == "Primary Tumor", ] # Define I-stage (include IA/IB) and IV-stage samples i_stage_ids <- tumor_meta$sample_id[grepl("Stage I", tumor_meta$pathologic_stage)] iv_stage_ids <- tumor_meta$sample_id[grepl("Stage IV", tumor_meta$pathologic_stage)] # Check sample counts (you need at least 3-5 samples per group for statistical power) cat("I-stage samples:", length(i_stage_ids), "\n") cat("IV-stage samples:", length(iv_stage_ids), "\n")
If either group has fewer than 3 samples, statistical tests won't have enough power to detect DEGs—you might need to pool adjacent stages (e.g., I+II vs III+IV) if individual stage groups are too small.
2. Confirm You're Using the Right Data Format for Your DEG Tool
FireBrowse's RNASeqV2 RSEM data comes in two flavors: raw counts and log2-transformed TPM/RSEM values. Using the wrong format with your analysis tool will lead to meaningless results:
- If you're using
DESeq2oredgeR, you need raw RSEM counts (not log-transformed). - If you're using
limma, log2-transformed values work, but you should usevoomif starting from raw counts, orlimma-trendif using pre-logged data.
Quick Check:
Look at your expression matrix values—if most values are between 0 and 15, it's likely log2(TPM+1) or log2(RSEM+1). If values are integers in the hundreds/thousands, it's raw counts.
3. Filter Out Low-Expression Genes
Low-expression genes (those with near-zero counts in most samples) add noise to your analysis and can mask true DEGs. Filter them out before running your test:
Example Filtering Code (R):
# For log2-transformed data (keep genes expressed >1 in at least 10% of samples) keep_genes <- rowSums(luad_expr > 1) >= 0.1 * ncol(luad_expr) luad_expr_filtered <- luad_expr[keep_genes, ] # For raw counts (keep genes with at least 10 counts in at least 3 samples) keep_genes <- rowSums(luad_raw_counts >= 10) >= 3 luad_counts_filtered <- luad_raw_counts[keep_genes, ]
4. Adjust Your Statistical Thresholds
You might be setting the bar too high:
- Try relaxing the FDR-adjusted p-value threshold from the default 0.05 to 0.1, especially if your sample sizes are small.
- Don't overlook genes with smaller logFC values—if you're filtering for |logFC| > 1, try lowering it to 0.5 to see if any genes come up.
Example with limma:
# After running eBayes, get results with relaxed thresholds deg_results <- topTable(fit2, coef = "StageIV - StageI", number = Inf, adjust = "fdr", p.value = 0.1, lfc = 0.5)
5. Account for Batch Effects & Confounders
FireBrowse samples are collected across multiple batches, which can introduce technical variation that swamps biological differences. Use tools like ComBat (from the sva package) to correct for batch effects before running DEG analysis:
Batch Correction Example:
library(sva) # Assume 'batch' is a column in luad_meta with batch IDs # Design matrix excluding stage (to avoid removing biological signal) mod <- model.matrix(~1, data = luad_meta) # Correct expression matrix for batch effects luad_expr_corrected <- ComBat(dat = luad_expr_filtered, batch = luad_meta$batch, mod = mod)
You should also include known confounders (like age, gender, or smoking status) in your design matrix if they're available in the metadata.
6. Try a Different DEG Method
If one tool isn't giving results, switch to another to rule out method-specific issues:
- If you used
DESeq2, tryedgeR(both use negative binomial models but have slight differences in normalization). - If you used
limma, tryDESeq2with raw counts.
内容的提问来源于stack exchange,提问作者Ankita Lawarde

