如何用Python实现含t-test与WMW test的贷款样本假设检验交叉表并导出CSV
Got it, let's walk through exactly how to build your cross-tabulation, run the required statistical tests, and export the results—using the libraries you already know (pandas, scipy) with clear, reproducible code.
1. Import Required Libraries
First, load all the tools we'll need:
import pandas as pd from scipy.stats import ttest_ind, mannwhitneyu import numpy as np
2. Load and Prepare Your Data
Assuming your Lending Club data is stored in a CSV file (adjust the path as needed). We'll also clean up any missing values in numerical columns to avoid errors in tests:
# Load data df = pd.read_csv("lending_club_loans.csv") # Filter numerical columns (we'll only run tests on these) num_cols = df.select_dtypes(include=[np.number]).columns.tolist() # Exclude the target variable 'dead' itself from stats calculations num_cols.remove('dead') # Drop rows with missing values in numerical columns (or impute if preferred) df_clean = df.dropna(subset=num_cols + ['dead']) # Split into sub-samples normal_loans = df_clean[df_clean['dead'] == 0] bad_loans = df_clean[df_clean['dead'] == 1]
3. Calculate Descriptive Statistics (Mean/Median)
We'll build a summary table with stats for the full sample, normal loans, and bad loans:
# Initialize a results dataframe stats_summary = pd.DataFrame(index=num_cols) # Add full sample stats stats_summary['Full Sample Mean'] = df_clean[num_cols].mean() stats_summary['Full Sample Median'] = df_clean[num_cols].median() # Add normal loan stats stats_summary['Normal Loans Mean'] = normal_loans[num_cols].mean() stats_summary['Normal Loans Median'] = normal_loans[num_cols].median() # Add bad loan stats stats_summary['Bad Loans Mean'] = bad_loans[num_cols].mean() stats_summary['Bad Loans Median'] = bad_loans[num_cols].median()
4. Run Statistical Tests
Now we'll add the t-test and Mann-Whitney U (WMW) test results, including p-values for significance:
# Helper function to run tests and return results def run_tests(col): # T-test (assuming equal variance; adjust equal_var=False if needed for Welch's test) t_stat, t_pval = ttest_ind(normal_loans[col], bad_loans[col], equal_var=True, nan_policy='omit') # Mann-Whitney U test (WMW) - two-sided alternative mw_stat, mw_pval = mannwhitneyu(normal_loans[col], bad_loans[col], alternative='two-sided') return pd.Series([t_stat, t_pval, mw_stat, mw_pval], index=['T-Statistic', 'T-Test P-Value', 'WMW Statistic', 'WMW P-Value']) # Apply the function to all numerical columns test_results = pd.DataFrame([run_tests(col) for col in num_cols], index=num_cols) # Combine descriptive stats with test results final_table = pd.concat([stats_summary, test_results], axis=1) # Optional: Add significance markers for quick interpretation final_table['T-Test Significance'] = np.where(final_table['T-Test P-Value'] < 0.01, '**', np.where(final_table['T-Test P-Value'] < 0.05, '*', '')) final_table['WMW Significance'] = np.where(final_table['WMW P-Value'] < 0.01, '**', np.where(final_table['WMW P-Value'] < 0.05, '*', ''))
5. Export to CSV
Save your final table to a CSV file for easy sharing or further analysis:
final_table.to_csv("lending_club_hypothesis_test_results.csv", index_label='Feature')
Optional: Interactive Table with Plotly
If you want to explore the results interactively (like you mentioned with plotly), here's a quick snippet to turn the final table into an interactive HTML table:
import plotly.graph_objects as go fig = go.Figure(data=[go.Table( header=dict(values=['Feature'] + list(final_table.columns), fill_color='paleturquoise', align='left'), cells=dict(values=[final_table.index] + [final_table[col] for col in final_table.columns], fill_color='lavender', align='left')) ]) fig.update_layout(title='Lending Club Loan Hypothesis Test Results') fig.write_html("interactive_test_results.html")
Quick Notes to Keep in Mind
- T-Test Assumptions: The t-test assumes normality and equal variances. If your data is skewed (common with loan data), the WMW test is more reliable. Use
equal_var=Falseif variances are unequal for Welch's t-test. - Missing Data: We dropped missing values here, but you could impute them (e.g., with median) if you want to retain more samples.
- Cross-Tabulation: If you want a formal cross-tab of categorical features against
dead, you can usepd.crosstab(df['categorical_col'], df['dead'])and add chi-squared tests if needed.
内容的提问来源于stack exchange,提问作者Mark Bartholomäus

