Stata 17中如何在汇总统计表格中添加变量正态性检验的P值?
Hey there! Great question—adding normality test p-values to your summary statistics table in Stata 17 is totally achievable, and I’ve got a few practical methods to help you pull this off smoothly.
Method 1: Use
estout for Publication-Ready Tables This is my go-to approach because it lets you generate clean, customizable tables that include both standard summary stats and normality test p-values. First, you’ll need the estout package (it’s free from Stata’s SSC repository):
* Install estout if you haven't already ssc install estout * Define your list of variables (replace with your actual variable names) local analysis_vars age income score * Clear existing estimates to start fresh estimates clear * Loop through each variable to collect stats and normality p-values foreach var of local analysis_vars { * Grab standard summary statistics estpost summarize `var' * Run Shapiro-Wilk normality test and save the p-value swilk `var' local sw_p = r(p) estadd scalar shapiro_wilk_p = `sw_p' * Optional: Add Kolmogorov-Smirnov test p-value (uses sample mean/sd as distribution params) ksmirnov `var' = normal(mean(`var'), sd(`var')) local ks_p = r(p) estadd scalar kolmogorov_smirnov_p = `ks_p' * Store the results for this variable estimates store `var' } * Generate the final table with your desired stats esttab *, cells("mean sd min max shapiro_wilk_p kolmogorov_smirnov_p") /// label /// title("Summary Statistics with Normality Test P-Values") /// format(%9.4f) /// replace
What this does:
- The loop processes each variable individually, collecting mean, standard deviation, min/max, and p-values from two common normality tests.
estadd scalarattaches the p-values to the summary stats so they show up in the final table.- The
esttabcommand lets you tweak formatting (like decimal places, labels, and title) to fit your needs.
Method 2: Manual Matrix Combination (No External Packages)
If you prefer not to install estout, you can manually compile stats and p-values into a matrix for display:
* Define your variables local analysis_vars age income score * Get basic summary stats and save them to a matrix tabstat `analysis_vars', stats(mean sd min max) save matrix basic_stats = r(StatTotal) * Create a matrix to hold normality p-values (2 rows for SW and KS tests) matrix norm_pvals = J(2, colsof(basic_stats), .) * Fill the p-value matrix forvalues i = 1/colsof(basic_stats) { local current_var = colnames(basic_stats)[`i'] * Shapiro-Wilk test swilk `current_var' matrix norm_pvals[1,`i'] = r(p) * Kolmogorov-Smirnov test ksmirnov `current_var' = normal(mean(`current_var'), sd(`current_var')) matrix norm_pvals[2,`i'] = r(p) } * Combine the basic stats and p-value matrices matrix full_stats = basic_stats \ norm_pvals * Add row/column labels for clarity matrix colnames full_stats = `:colnames(basic_stats)' matrix rownames full_stats = "Mean" "Standard Deviation" "Minimum" "Maximum" "Shapiro-Wilk P" "Kolmogorov-Smirnov P" * Display the final matrix matrix list full_stats, format(%9.4f)
Method 3: Quick In-Line Output (For Quick Checks)
If you just want to view stats and p-values side-by-side without a formal table, a simple loop will do:
local analysis_vars age income score foreach var of local analysis_vars { di "=====================================" di "Summary Statistics for: `var'" di "=====================================" summarize `var' di "" swilk `var' di "Shapiro-Wilk Normality Test P-Value: " string(r(p), "%9.4f") ksmirnov `var' = normal(mean(`var'), sd(`var')) di "Kolmogorov-Smirnov Normality Test P-Value: " string(r(p), "%9.4f") di "" }
A Quick Note on Normality Tests:
- Shapiro-Wilk: Best suited for small to medium-sized samples (typically n ≤ 2000; Stata may restrict it for larger samples).
- Kolmogorov-Smirnov: Works for larger samples, but note that we’re using the sample mean and standard deviation to define the normal distribution here (this is the "Lilliefors" variant in practice).
内容的提问来源于stack exchange,提问作者Patudb
相关产品推荐
相关产品推荐

