You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在R中批量运行t.tests比较成对列值的实现方法

Batch Running t-tests for Corresponding A#U/C#U Column Pairs

Hey there! Let's solve this problem step by step—batch processing those t-tests doesn't have to be tricky. I'll walk you through identifying the column pairs, running the tests, and organizing the results clearly.

Step 1: Identify Matching Column Pairs

First, we need to map each A#U column to its corresponding C#U column. We can do this by extracting the numeric part from the A columns and checking if the matching C column exists in your dataframe.

import pandas as pd
from scipy.stats import ttest_ind, ttest_rel

# Assume your dataframe is named df
df = pd.read_csv("your_data.csv")  # Or whatever way you load your data

# Get all A#U columns
a_columns = [col for col in df.columns if col.startswith("A") and col.endswith("U")]

# Create list of (A_col, C_col) pairs
column_pairs = []
for a_col in a_columns:
    # Extract the numeric part (e.g., "1" from "A1U")
    number_part = a_col[1:-1]  # Slice between "A" and "U"
    c_col = f"C{number_part}U"
    
    if c_col in df.columns:
        column_pairs.append((a_col, c_col))
    else:
        print(f"Warning: No matching column {c_col} found for {a_col}")

Step 2: Run t-tests for Each Pair

Next, we'll loop through each pair, run the appropriate t-test, and collect the results. I'll cover both independent and paired t-tests since which one you use depends on your study design.

Option A: Independent Samples t-test (Welch's)

Use this if the A#U and C#U columns are from separate groups:

results = []
for a_col, c_col in column_pairs:
    # Drop missing values to avoid errors
    group_a = df[a_col].dropna()
    group_c = df[c_col].dropna()
    
    # Run Welch's t-test (equal_var=False accounts for unequal variances)
    t_stat, p_val = ttest_ind(group_a, group_c, equal_var=False)
    
    # Store results
    results.append({
        "Column Pair": f"{a_col} vs {c_col}",
        "t-statistic": round(t_stat, 4),
        "p-value": round(p_val, 4),
        "Degrees of Freedom": round(ttest_ind(group_a, group_c, equal_var=False).df, 2)
    })

# Convert results to a readable dataframe
results_df = pd.DataFrame(results)
print(results_df)

Option B: Paired t-test

Use this if each row in A#U and C#U corresponds to the same subject/observation:

results = []
for a_col, c_col in column_pairs:
    # Drop rows where either column has a missing value (critical for paired tests)
    paired_data = df[[a_col, c_col]].dropna()
    group_a = paired_data[a_col]
    group_c = paired_data[c_col]
    
    # Run paired t-test
    t_stat, p_val = ttest_rel(group_a, group_c)
    
    # Store results
    results.append({
        "Column Pair": f"{a_col} vs {c_col}",
        "t-statistic": round(t_stat, 4),
        "p-value": round(p_val, 4),
        "Degrees of Freedom": len(paired_data) - 1
    })

results_df = pd.DataFrame(results)
print(results_df)

Why Your Previous Attempts Might Have Failed

  • Incorrect Column Matching: If your loop/apply wasn't properly extracting the numeric part to match A and C columns, you might have been comparing the wrong pairs.
  • Missing Values: Scipy's t-test functions will throw errors if there are NaNs in the data—always drop or impute missing values first.
  • Wrong Test Type: Using an independent test for paired data (or vice versa) can lead to incorrect results, and might have caused unexpected behavior in your code.

Bonus: Quick One-Liner Alternative (Using List Comprehension)

If you prefer a more concise approach, you can generate the results with a list comprehension:

# For independent tests
results = [
    {
        "Pair": f"{a} vs {c}",
        "t": round(ttest_ind(df[a].dropna(), df[c].dropna(), equal_var=False)[0],4),
        "p": round(ttest_ind(df[a].dropna(), df[c].dropna(), equal_var=False)[1],4)
    }
    for a, c in column_pairs
]

内容的提问来源于stack exchange,提问作者melbez

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 07:26:24