在R中批量运行t.tests比较成对列值的实现方法
Hey there! Let's solve this problem step by step—batch processing those t-tests doesn't have to be tricky. I'll walk you through identifying the column pairs, running the tests, and organizing the results clearly.
Step 1: Identify Matching Column Pairs
First, we need to map each A#U column to its corresponding C#U column. We can do this by extracting the numeric part from the A columns and checking if the matching C column exists in your dataframe.
import pandas as pd from scipy.stats import ttest_ind, ttest_rel # Assume your dataframe is named df df = pd.read_csv("your_data.csv") # Or whatever way you load your data # Get all A#U columns a_columns = [col for col in df.columns if col.startswith("A") and col.endswith("U")] # Create list of (A_col, C_col) pairs column_pairs = [] for a_col in a_columns: # Extract the numeric part (e.g., "1" from "A1U") number_part = a_col[1:-1] # Slice between "A" and "U" c_col = f"C{number_part}U" if c_col in df.columns: column_pairs.append((a_col, c_col)) else: print(f"Warning: No matching column {c_col} found for {a_col}")
Step 2: Run t-tests for Each Pair
Next, we'll loop through each pair, run the appropriate t-test, and collect the results. I'll cover both independent and paired t-tests since which one you use depends on your study design.
Option A: Independent Samples t-test (Welch's)
Use this if the A#U and C#U columns are from separate groups:
results = [] for a_col, c_col in column_pairs: # Drop missing values to avoid errors group_a = df[a_col].dropna() group_c = df[c_col].dropna() # Run Welch's t-test (equal_var=False accounts for unequal variances) t_stat, p_val = ttest_ind(group_a, group_c, equal_var=False) # Store results results.append({ "Column Pair": f"{a_col} vs {c_col}", "t-statistic": round(t_stat, 4), "p-value": round(p_val, 4), "Degrees of Freedom": round(ttest_ind(group_a, group_c, equal_var=False).df, 2) }) # Convert results to a readable dataframe results_df = pd.DataFrame(results) print(results_df)
Option B: Paired t-test
Use this if each row in A#U and C#U corresponds to the same subject/observation:
results = [] for a_col, c_col in column_pairs: # Drop rows where either column has a missing value (critical for paired tests) paired_data = df[[a_col, c_col]].dropna() group_a = paired_data[a_col] group_c = paired_data[c_col] # Run paired t-test t_stat, p_val = ttest_rel(group_a, group_c) # Store results results.append({ "Column Pair": f"{a_col} vs {c_col}", "t-statistic": round(t_stat, 4), "p-value": round(p_val, 4), "Degrees of Freedom": len(paired_data) - 1 }) results_df = pd.DataFrame(results) print(results_df)
Why Your Previous Attempts Might Have Failed
- Incorrect Column Matching: If your loop/apply wasn't properly extracting the numeric part to match
AandCcolumns, you might have been comparing the wrong pairs. - Missing Values: Scipy's t-test functions will throw errors if there are NaNs in the data—always drop or impute missing values first.
- Wrong Test Type: Using an independent test for paired data (or vice versa) can lead to incorrect results, and might have caused unexpected behavior in your code.
Bonus: Quick One-Liner Alternative (Using List Comprehension)
If you prefer a more concise approach, you can generate the results with a list comprehension:
# For independent tests results = [ { "Pair": f"{a} vs {c}", "t": round(ttest_ind(df[a].dropna(), df[c].dropna(), equal_var=False)[0],4), "p": round(ttest_ind(df[a].dropna(), df[c].dropna(), equal_var=False)[1],4) } for a, c in column_pairs ]
内容的提问来源于stack exchange,提问作者melbez

