You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于两个Pandas DataFrame的T检验技术咨询

Running Independent Samples T-Tests Between Normal and Tumor Gene Expression DataFrames

Alright, let's walk through how to run t-tests between your Normal (df) and Tumor (df2) DataFrames step by step. First, let's confirm the setup: each row represents a gene, and each column is a sample from either the Normal or Tumor group—perfect, that's exactly the structure we need for per-gene comparisons.

Step 1: Align Your Data

First, make sure both DataFrames have the same set of genes (in the same order) to avoid mismatched comparisons. If there are any genes missing in one dataset, we'll filter to keep only the common ones:

import pandas as pd
from scipy.stats import ttest_ind

# Get genes present in both DataFrames
common_genes = df.index.intersection(df2.index)
# Filter both DataFrames to keep only common genes
df_normal = df.loc[common_genes]
df_tumor = df2.loc[common_genes]

Step 2: Choose the Right T-Test

Since your samples are from two independent groups (Normal vs. Tumor), we'll use the independent samples t-test. For most gene expression data, we use Welch's t-test (instead of the standard Student's t-test) because variance between groups is often unequal. Welch's t-test adjusts for this automatically.

Step 3: Run Per-Gene T-Tests

We can compute t-statistics and p-values for every gene efficiently using vectorized operations (no slow loops needed):

# Calculate t-statistics and p-values across all genes (axis=1 means row-wise)
t_stats, p_values = ttest_ind(
    df_normal.values, 
    df_tumor.values, 
    axis=1, 
    equal_var=False,  # Enables Welch's t-test
    nan_policy='omit'  # Ignore any missing values if present
)

# Build a results DataFrame for easy analysis
ttest_results = pd.DataFrame({
    'mean_normal': df_normal.mean(axis=1),
    'mean_tumor': df_tumor.mean(axis=1),
    't_statistic': t_stats,
    'p_value': p_values
}, index=common_genes)

Step 4: Adjust for Multiple Comparisons

When testing hundreds of genes, you need to correct p-values to avoid false positives. The Bonferroni correction is a simple (though conservative) method:

# Apply Bonferroni correction
ttest_results['adjusted_p_value'] = ttest_results['p_value'] * len(ttest_results)
# Ensure adjusted p-values don't exceed 1 (a quirk of the correction)
ttest_results['adjusted_p_value'] = ttest_results['adjusted_p_value'].clip(upper=1)

Key Notes & Edge Cases

  • Paired Samples? If your data is paired (e.g., same patient's Normal and Tumor sample), use ttest_rel instead of ttest_ind—just make sure the columns in df and df2 are in matching order.
  • Small Sample Sizes? For small sample groups (n < 30), you should check if your data follows a normal distribution first (e.g., Shapiro-Wilk test). Welch's t-test is more robust to non-normality than Student's, but extreme deviations can still affect results.
  • Outliers? Gene expression data often has outliers—consider trimming extreme values or using a non-parametric test (like the Mann-Whitney U test) if outliers are skewing your results.

Example Output

Your ttest_results DataFrame will look something like this:

genemean_normalmean_tumort_statisticp_valueadjusted_p_value
gene01.8318.67-2.140.0354.86
gene134.7257.21-1.980.0496.86
..................

You can filter for significantly differentially expressed genes using:

# Genes with adjusted p-value < 0.05 and mean fold change > 2 (example threshold)
significant_genes = ttest_results[
    (ttest_results['adjusted_p_value'] < 0.05) & 
    (abs(ttest_results['mean_tumor'] / ttest_results['mean_normal']) > 2)
]

内容的提问来源于stack exchange,提问作者Rujun Guan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 03:16:53