基于超几何分布的基因表达重叠分析:RNA-seq数据处理难题
Hey there, let's work through how to use the hypergeometric distribution to test the significance of overlapping DE genes between your D1 and D2 groups. From your description, you've already filtered out genes that are DE in the Control group, which is a smart step to focus on treatment-specific effects.
Step 1: Nail Down the Hypergeometric Parameters
First, let's define the four key values you need for the test—these come straight from your dataset:
- N: The total number of genes in your background set. This should be all genes that are not DE in the Control group. Since your full transcriptome has 15000 genes, if Control has
CDE genes, thenN = 15000 - C. - K: The number of DE genes in the D1 group (after excluding Control DEs). Let's call this
D1_induced_DE. - n: The number of DE genes in the D2 group (after excluding Control DEs). Let's name this
D2_induced_DE. - k: The number of genes that are DE in both D1 and D2 (and not DE in Control)—this is your overlapping gene set.
Step 2: Calculate the Statistical Significance
R is the go-to tool for this kind of test, and the phyper() function makes it straightforward. We'll run a one-tailed test to see if the overlap is larger than what you'd expect by random chance:
# Compute the p-value p_value <- phyper(q = k - 1, m = K, n = N - K, k = n, lower.tail = FALSE)
Let's break down what each argument does:
q = k - 1: We subtract 1 because we want the probability of observing at least k overlapping genes (not just exactly k).m = K: The number of "successes" in the population—these are the D1-induced DE genes.n = N - K: The number of "non-successes" in the population—genes that aren't DE in D1 or Control.k = n: The size of our sample—this is the D2-induced DE genes.lower.tail = FALSE: Tells the function to return the upper tail probability (the chance of getting q or more overlaps).
Step 3: Interpret the Result
- A small p-value (usually ≤ 0.05) means your overlap is statistically significant—so it's unlikely the shared DE genes are just a random coincidence.
- A large p-value suggests the overlap could be due to chance, so there's no strong evidence that D1 and D2 are regulating the same set of genes (beyond random variation).
Quick Example to Illustrate
Let's plug in some hypothetical numbers to make this concrete:
- Control has 200 DE genes →
N = 15000 - 200 = 14800 - D1 has 600 induced DE genes →
K = 600 - D2 has 500 induced DE genes →
n = 500 - Overlap between D1 and D2 is 250 →
k = 250
Running the code:
p_value <- phyper(q = 249, m = 600, n = 14800 - 600, k = 500, lower.tail = FALSE) print(p_value)
This would give you the p-value for how significant that 250-gene overlap is.
Important Checks
- Double-check your background set: Make sure
Nonly includes genes that aren't DE in Control—this keeps your test focused on treatment-specific effects, not pre-existing differences in Control. - If you were running multiple overlap tests, you'd want to adjust for multiple comparisons (like Bonferroni), but since this is a single test, it's probably not necessary here.
内容的提问来源于stack exchange,提问作者norty89

