基于两个Pandas DataFrame的T检验技术咨询
Alright, let's walk through how to run t-tests between your Normal (df) and Tumor (df2) DataFrames step by step. First, let's confirm the setup: each row represents a gene, and each column is a sample from either the Normal or Tumor group—perfect, that's exactly the structure we need for per-gene comparisons.
Step 1: Align Your Data
First, make sure both DataFrames have the same set of genes (in the same order) to avoid mismatched comparisons. If there are any genes missing in one dataset, we'll filter to keep only the common ones:
import pandas as pd from scipy.stats import ttest_ind # Get genes present in both DataFrames common_genes = df.index.intersection(df2.index) # Filter both DataFrames to keep only common genes df_normal = df.loc[common_genes] df_tumor = df2.loc[common_genes]
Step 2: Choose the Right T-Test
Since your samples are from two independent groups (Normal vs. Tumor), we'll use the independent samples t-test. For most gene expression data, we use Welch's t-test (instead of the standard Student's t-test) because variance between groups is often unequal. Welch's t-test adjusts for this automatically.
Step 3: Run Per-Gene T-Tests
We can compute t-statistics and p-values for every gene efficiently using vectorized operations (no slow loops needed):
# Calculate t-statistics and p-values across all genes (axis=1 means row-wise) t_stats, p_values = ttest_ind( df_normal.values, df_tumor.values, axis=1, equal_var=False, # Enables Welch's t-test nan_policy='omit' # Ignore any missing values if present ) # Build a results DataFrame for easy analysis ttest_results = pd.DataFrame({ 'mean_normal': df_normal.mean(axis=1), 'mean_tumor': df_tumor.mean(axis=1), 't_statistic': t_stats, 'p_value': p_values }, index=common_genes)
Step 4: Adjust for Multiple Comparisons
When testing hundreds of genes, you need to correct p-values to avoid false positives. The Bonferroni correction is a simple (though conservative) method:
# Apply Bonferroni correction ttest_results['adjusted_p_value'] = ttest_results['p_value'] * len(ttest_results) # Ensure adjusted p-values don't exceed 1 (a quirk of the correction) ttest_results['adjusted_p_value'] = ttest_results['adjusted_p_value'].clip(upper=1)
Key Notes & Edge Cases
- Paired Samples? If your data is paired (e.g., same patient's Normal and Tumor sample), use
ttest_relinstead ofttest_ind—just make sure the columns indfanddf2are in matching order. - Small Sample Sizes? For small sample groups (n < 30), you should check if your data follows a normal distribution first (e.g., Shapiro-Wilk test). Welch's t-test is more robust to non-normality than Student's, but extreme deviations can still affect results.
- Outliers? Gene expression data often has outliers—consider trimming extreme values or using a non-parametric test (like the Mann-Whitney U test) if outliers are skewing your results.
Example Output
Your ttest_results DataFrame will look something like this:
| gene | mean_normal | mean_tumor | t_statistic | p_value | adjusted_p_value |
|---|---|---|---|---|---|
| gene0 | 1.83 | 18.67 | -2.14 | 0.035 | 4.86 |
| gene1 | 34.72 | 57.21 | -1.98 | 0.049 | 6.86 |
| ... | ... | ... | ... | ... | ... |
You can filter for significantly differentially expressed genes using:
# Genes with adjusted p-value < 0.05 and mean fold change > 2 (example threshold) significant_genes = ttest_results[ (ttest_results['adjusted_p_value'] < 0.05) & (abs(ttest_results['mean_tumor'] / ttest_results['mean_normal']) > 2) ]
内容的提问来源于stack exchange,提问作者Rujun Guan

