You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python/Pandas实现正态分布转换及类似Stata的ladder功能?

Replicating Stata's ladder/gladder in Pandas/Python

Great question! I’ve missed having Stata’s handy ladder and gladder tools when working in Python too, but you can absolutely build a similar workflow using Pandas alongside libraries like scipy.stats and matplotlib/seaborn. Here’s how to do it:

Core Idea

Stata’s ladder tests a suite of common data transformations (log, square root, reciprocal, etc.) and reports statistics that measure how close each transformed dataset is to a normal distribution. gladder takes this a step further with visualizations (histograms, Q-Q plots) for each transformation. We’ll replicate both the numerical assessment and the visual checks.

Step 1: Define Common Transformations

First, let’s list out the transformations we want to test—these match what ladder uses, plus a couple of more robust options for tricky data:

import pandas as pd
import numpy as np
import scipy.stats as stats
import matplotlib.pyplot as plt
import seaborn as sns

def apply_transformations(series):
    transforms = {
        "Original": series,
        "Log (x)": np.log(series),
        "Log (x+1)": np.log1p(series),  # For non-positive data
        "Square Root": np.sqrt(series),
        "Reciprocal": 1 / series,
        "Square": series ** 2,
        "Cube": series ** 3
    }
    # Add Box-Cox transformation (requires all positive data)
    if (series > 0).all():
        boxcox_data, _ = stats.boxcox(series)
        transforms["Box-Cox"] = boxcox_data
    # Add Yeo-Johnson (works with negative/zero data)
    yeojohnson_data, _ = stats.yeojohnson(series)
    transforms["Yeo-Johnson"] = yeojohnson_data
    
    return transforms

Step 2: Calculate Normality Metrics

Next, we’ll compute key statistics to evaluate normality for each transformed dataset. We’ll use:

  • Shapiro-Wilk Test: Ideal for small-to-medium datasets (p-value > 0.05 suggests normality)
  • Skewness & Kurtosis: Values close to 0 indicate a normal-like distribution
  • Kolmogorov-Smirnov Test: Compares the dataset to a theoretical normal distribution
def calculate_normality_metrics(transformed_data):
    results = []
    for name, data in transformed_data.items():
        # Drop NaNs from transformations that might introduce them
        clean_data = data.dropna()
        if len(clean_data) < 3:  # Skip if too few valid points
            continue
        
        shapiro_stat, shapiro_p = stats.shapiro(clean_data)
        skewness = stats.skew(clean_data)
        kurtosis = stats.kurtosis(clean_data)
        ks_stat, ks_p = stats.kstest(clean_data, 'norm', args=(clean_data.mean(), clean_data.std()))
        
        results.append({
            "Transformation": name,
            "Shapiro-Wilk p-value": round(shapiro_p, 4),
            "Skewness": round(skewness, 4),
            "Kurtosis": round(kurtosis, 4),
            "KS p-value": round(ks_p, 4)
        })
    
    return pd.DataFrame(results).sort_values("Shapiro-Wilk p-value", ascending=False)

Step 3: Generate the "Ladder" Table

Let’s put it all together with a sample skewed dataset (the kind where transformations are most useful):

# Create a sample right-skewed dataset
np.random.seed(42)
data = pd.Series(np.random.exponential(scale=10, size=200), name="Value")

# Apply transformations and get normality metrics
transformed_data = apply_transformations(data)
normality_table = calculate_normality_metrics(transformed_data)

print(normality_table)

This table will show you which transformations make your data most normal—look for the highest Shapiro-Wilk/KS p-values and skewness/kurtosis values closest to 0.

Step 4: Build the "Gladder" Visualization

Finally, let’s replicate the visual checks with subplots for each transformation’s histogram (overlaid with a normal curve) and Q-Q plot:

def plot_gladder(transformed_data):
    num_plots = len(transformed_data)
    fig, axes = plt.subplots(num_plots, 2, figsize=(12, 4*num_plots))
    plt.subplots_adjust(hspace=0.5)
    
    for i, (name, data) in enumerate(transformed_data.items()):
        clean_data = data.dropna()
        if len(clean_data) < 3:
            continue
        
        # Histogram with normal curve overlay
        ax_hist = axes[i, 0]
        sns.histplot(clean_data, kde=False, ax=ax_hist, bins=15)
        # Plot matching normal curve
        mu, sigma = clean_data.mean(), clean_data.std()
        x = np.linspace(mu - 3*sigma, mu + 3*sigma, 100)
        ax_hist.plot(x, stats.norm.pdf(x, mu, sigma)*len(clean_data)*sigma*0.8, 'r--')
        ax_hist.set_title(f"{name} - Histogram")
        
        # Q-Q Plot
        ax_qq = axes[i, 1]
        stats.probplot(clean_data, plot=ax_qq)
        ax_qq.set_title(f"{name} - Q-Q Plot")
    
    plt.show()

plot_gladder(transformed_data)

This will generate a grid of plots where you can visually inspect which transformations align best with a normal distribution—points on the Q-Q plot should lie close to the diagonal line for well-behaved data.

Bonus Tips

  • If your data has negative values, prioritize Log(x+1) or Yeo-Johnson transformations since log/square root won’t work on negatives.
  • For large datasets, the Shapiro-Wilk test will often reject normality even for minor deviations—stick to visual checks or the KS test in those cases.
  • You can customize the list of transformations to include others that make sense for your data (e.g., log base 10, cube root).

内容的提问来源于stack exchange,提问作者col. slade

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 08:56:31