如何用Python/Pandas实现正态分布转换及类似Stata的ladder功能?
ladder/gladder in Pandas/Python Great question! I’ve missed having Stata’s handy ladder and gladder tools when working in Python too, but you can absolutely build a similar workflow using Pandas alongside libraries like scipy.stats and matplotlib/seaborn. Here’s how to do it:
Core Idea
Stata’s ladder tests a suite of common data transformations (log, square root, reciprocal, etc.) and reports statistics that measure how close each transformed dataset is to a normal distribution. gladder takes this a step further with visualizations (histograms, Q-Q plots) for each transformation. We’ll replicate both the numerical assessment and the visual checks.
Step 1: Define Common Transformations
First, let’s list out the transformations we want to test—these match what ladder uses, plus a couple of more robust options for tricky data:
import pandas as pd import numpy as np import scipy.stats as stats import matplotlib.pyplot as plt import seaborn as sns def apply_transformations(series): transforms = { "Original": series, "Log (x)": np.log(series), "Log (x+1)": np.log1p(series), # For non-positive data "Square Root": np.sqrt(series), "Reciprocal": 1 / series, "Square": series ** 2, "Cube": series ** 3 } # Add Box-Cox transformation (requires all positive data) if (series > 0).all(): boxcox_data, _ = stats.boxcox(series) transforms["Box-Cox"] = boxcox_data # Add Yeo-Johnson (works with negative/zero data) yeojohnson_data, _ = stats.yeojohnson(series) transforms["Yeo-Johnson"] = yeojohnson_data return transforms
Step 2: Calculate Normality Metrics
Next, we’ll compute key statistics to evaluate normality for each transformed dataset. We’ll use:
- Shapiro-Wilk Test: Ideal for small-to-medium datasets (p-value > 0.05 suggests normality)
- Skewness & Kurtosis: Values close to 0 indicate a normal-like distribution
- Kolmogorov-Smirnov Test: Compares the dataset to a theoretical normal distribution
def calculate_normality_metrics(transformed_data): results = [] for name, data in transformed_data.items(): # Drop NaNs from transformations that might introduce them clean_data = data.dropna() if len(clean_data) < 3: # Skip if too few valid points continue shapiro_stat, shapiro_p = stats.shapiro(clean_data) skewness = stats.skew(clean_data) kurtosis = stats.kurtosis(clean_data) ks_stat, ks_p = stats.kstest(clean_data, 'norm', args=(clean_data.mean(), clean_data.std())) results.append({ "Transformation": name, "Shapiro-Wilk p-value": round(shapiro_p, 4), "Skewness": round(skewness, 4), "Kurtosis": round(kurtosis, 4), "KS p-value": round(ks_p, 4) }) return pd.DataFrame(results).sort_values("Shapiro-Wilk p-value", ascending=False)
Step 3: Generate the "Ladder" Table
Let’s put it all together with a sample skewed dataset (the kind where transformations are most useful):
# Create a sample right-skewed dataset np.random.seed(42) data = pd.Series(np.random.exponential(scale=10, size=200), name="Value") # Apply transformations and get normality metrics transformed_data = apply_transformations(data) normality_table = calculate_normality_metrics(transformed_data) print(normality_table)
This table will show you which transformations make your data most normal—look for the highest Shapiro-Wilk/KS p-values and skewness/kurtosis values closest to 0.
Step 4: Build the "Gladder" Visualization
Finally, let’s replicate the visual checks with subplots for each transformation’s histogram (overlaid with a normal curve) and Q-Q plot:
def plot_gladder(transformed_data): num_plots = len(transformed_data) fig, axes = plt.subplots(num_plots, 2, figsize=(12, 4*num_plots)) plt.subplots_adjust(hspace=0.5) for i, (name, data) in enumerate(transformed_data.items()): clean_data = data.dropna() if len(clean_data) < 3: continue # Histogram with normal curve overlay ax_hist = axes[i, 0] sns.histplot(clean_data, kde=False, ax=ax_hist, bins=15) # Plot matching normal curve mu, sigma = clean_data.mean(), clean_data.std() x = np.linspace(mu - 3*sigma, mu + 3*sigma, 100) ax_hist.plot(x, stats.norm.pdf(x, mu, sigma)*len(clean_data)*sigma*0.8, 'r--') ax_hist.set_title(f"{name} - Histogram") # Q-Q Plot ax_qq = axes[i, 1] stats.probplot(clean_data, plot=ax_qq) ax_qq.set_title(f"{name} - Q-Q Plot") plt.show() plot_gladder(transformed_data)
This will generate a grid of plots where you can visually inspect which transformations align best with a normal distribution—points on the Q-Q plot should lie close to the diagonal line for well-behaved data.
Bonus Tips
- If your data has negative values, prioritize
Log(x+1)orYeo-Johnsontransformations since log/square root won’t work on negatives. - For large datasets, the Shapiro-Wilk test will often reject normality even for minor deviations—stick to visual checks or the KS test in those cases.
- You can customize the list of transformations to include others that make sense for your data (e.g., log base 10, cube root).
内容的提问来源于stack exchange,提问作者col. slade

