You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python实现含t-test与WMW test的贷款样本假设检验交叉表并导出CSV

Step-by-Step Python Implementation for Lending Club Loan Hypothesis Testing

Got it, let's walk through exactly how to build your cross-tabulation, run the required statistical tests, and export the results—using the libraries you already know (pandas, scipy) with clear, reproducible code.

1. Import Required Libraries

First, load all the tools we'll need:

import pandas as pd
from scipy.stats import ttest_ind, mannwhitneyu
import numpy as np

2. Load and Prepare Your Data

Assuming your Lending Club data is stored in a CSV file (adjust the path as needed). We'll also clean up any missing values in numerical columns to avoid errors in tests:

# Load data
df = pd.read_csv("lending_club_loans.csv")

# Filter numerical columns (we'll only run tests on these)
num_cols = df.select_dtypes(include=[np.number]).columns.tolist()
# Exclude the target variable 'dead' itself from stats calculations
num_cols.remove('dead')

# Drop rows with missing values in numerical columns (or impute if preferred)
df_clean = df.dropna(subset=num_cols + ['dead'])

# Split into sub-samples
normal_loans = df_clean[df_clean['dead'] == 0]
bad_loans = df_clean[df_clean['dead'] == 1]

3. Calculate Descriptive Statistics (Mean/Median)

We'll build a summary table with stats for the full sample, normal loans, and bad loans:

# Initialize a results dataframe
stats_summary = pd.DataFrame(index=num_cols)

# Add full sample stats
stats_summary['Full Sample Mean'] = df_clean[num_cols].mean()
stats_summary['Full Sample Median'] = df_clean[num_cols].median()

# Add normal loan stats
stats_summary['Normal Loans Mean'] = normal_loans[num_cols].mean()
stats_summary['Normal Loans Median'] = normal_loans[num_cols].median()

# Add bad loan stats
stats_summary['Bad Loans Mean'] = bad_loans[num_cols].mean()
stats_summary['Bad Loans Median'] = bad_loans[num_cols].median()

4. Run Statistical Tests

Now we'll add the t-test and Mann-Whitney U (WMW) test results, including p-values for significance:

# Helper function to run tests and return results
def run_tests(col):
    # T-test (assuming equal variance; adjust equal_var=False if needed for Welch's test)
    t_stat, t_pval = ttest_ind(normal_loans[col], bad_loans[col], equal_var=True, nan_policy='omit')
    # Mann-Whitney U test (WMW) - two-sided alternative
    mw_stat, mw_pval = mannwhitneyu(normal_loans[col], bad_loans[col], alternative='two-sided')
    return pd.Series([t_stat, t_pval, mw_stat, mw_pval], 
                     index=['T-Statistic', 'T-Test P-Value', 'WMW Statistic', 'WMW P-Value'])

# Apply the function to all numerical columns
test_results = pd.DataFrame([run_tests(col) for col in num_cols], index=num_cols)

# Combine descriptive stats with test results
final_table = pd.concat([stats_summary, test_results], axis=1)

# Optional: Add significance markers for quick interpretation
final_table['T-Test Significance'] = np.where(final_table['T-Test P-Value'] < 0.01, '**', 
                                              np.where(final_table['T-Test P-Value'] < 0.05, '*', ''))
final_table['WMW Significance'] = np.where(final_table['WMW P-Value'] < 0.01, '**', 
                                           np.where(final_table['WMW P-Value'] < 0.05, '*', ''))

5. Export to CSV

Save your final table to a CSV file for easy sharing or further analysis:

final_table.to_csv("lending_club_hypothesis_test_results.csv", index_label='Feature')

Optional: Interactive Table with Plotly

If you want to explore the results interactively (like you mentioned with plotly), here's a quick snippet to turn the final table into an interactive HTML table:

import plotly.graph_objects as go

fig = go.Figure(data=[go.Table(
    header=dict(values=['Feature'] + list(final_table.columns),
                fill_color='paleturquoise',
                align='left'),
    cells=dict(values=[final_table.index] + [final_table[col] for col in final_table.columns],
               fill_color='lavender',
               align='left'))
])

fig.update_layout(title='Lending Club Loan Hypothesis Test Results')
fig.write_html("interactive_test_results.html")

Quick Notes to Keep in Mind

  • T-Test Assumptions: The t-test assumes normality and equal variances. If your data is skewed (common with loan data), the WMW test is more reliable. Use equal_var=False if variances are unequal for Welch's t-test.
  • Missing Data: We dropped missing values here, but you could impute them (e.g., with median) if you want to retain more samples.
  • Cross-Tabulation: If you want a formal cross-tab of categorical features against dead, you can use pd.crosstab(df['categorical_col'], df['dead']) and add chi-squared tests if needed.

内容的提问来源于stack exchange,提问作者Mark Bartholomäus

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.13 07:28:27