You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Stata 17中如何在汇总统计表格中添加变量正态性检验的P值?

Hey there! Great question—adding normality test p-values to your summary statistics table in Stata 17 is totally achievable, and I’ve got a few practical methods to help you pull this off smoothly.

Method 1: Use estout for Publication-Ready Tables

This is my go-to approach because it lets you generate clean, customizable tables that include both standard summary stats and normality test p-values. First, you’ll need the estout package (it’s free from Stata’s SSC repository):

* Install estout if you haven't already
ssc install estout

* Define your list of variables (replace with your actual variable names)
local analysis_vars age income score

* Clear existing estimates to start fresh
estimates clear

* Loop through each variable to collect stats and normality p-values
foreach var of local analysis_vars {
    * Grab standard summary statistics
    estpost summarize `var'
    
    * Run Shapiro-Wilk normality test and save the p-value
    swilk `var'
    local sw_p = r(p)
    estadd scalar shapiro_wilk_p = `sw_p'
    
    * Optional: Add Kolmogorov-Smirnov test p-value (uses sample mean/sd as distribution params)
    ksmirnov `var' = normal(mean(`var'), sd(`var'))
    local ks_p = r(p)
    estadd scalar kolmogorov_smirnov_p = `ks_p'
    
    * Store the results for this variable
    estimates store `var'
}

* Generate the final table with your desired stats
esttab *, cells("mean sd min max shapiro_wilk_p kolmogorov_smirnov_p") ///
        label ///
        title("Summary Statistics with Normality Test P-Values") ///
        format(%9.4f) ///
        replace

What this does:

  • The loop processes each variable individually, collecting mean, standard deviation, min/max, and p-values from two common normality tests.
  • estadd scalar attaches the p-values to the summary stats so they show up in the final table.
  • The esttab command lets you tweak formatting (like decimal places, labels, and title) to fit your needs.
Method 2: Manual Matrix Combination (No External Packages)

If you prefer not to install estout, you can manually compile stats and p-values into a matrix for display:

* Define your variables
local analysis_vars age income score

* Get basic summary stats and save them to a matrix
tabstat `analysis_vars', stats(mean sd min max) save
matrix basic_stats = r(StatTotal)

* Create a matrix to hold normality p-values (2 rows for SW and KS tests)
matrix norm_pvals = J(2, colsof(basic_stats), .)

* Fill the p-value matrix
forvalues i = 1/colsof(basic_stats) {
    local current_var = colnames(basic_stats)[`i']
    
    * Shapiro-Wilk test
    swilk `current_var'
    matrix norm_pvals[1,`i'] = r(p)
    
    * Kolmogorov-Smirnov test
    ksmirnov `current_var' = normal(mean(`current_var'), sd(`current_var'))
    matrix norm_pvals[2,`i'] = r(p)
}

* Combine the basic stats and p-value matrices
matrix full_stats = basic_stats \ norm_pvals

* Add row/column labels for clarity
matrix colnames full_stats = `:colnames(basic_stats)'
matrix rownames full_stats = "Mean" "Standard Deviation" "Minimum" "Maximum" "Shapiro-Wilk P" "Kolmogorov-Smirnov P"

* Display the final matrix
matrix list full_stats, format(%9.4f)
Method 3: Quick In-Line Output (For Quick Checks)

If you just want to view stats and p-values side-by-side without a formal table, a simple loop will do:

local analysis_vars age income score

foreach var of local analysis_vars {
    di "====================================="
    di "Summary Statistics for: `var'"
    di "====================================="
    summarize `var'
    di ""
    swilk `var'
    di "Shapiro-Wilk Normality Test P-Value: " string(r(p), "%9.4f")
    ksmirnov `var' = normal(mean(`var'), sd(`var'))
    di "Kolmogorov-Smirnov Normality Test P-Value: " string(r(p), "%9.4f")
    di ""
}

A Quick Note on Normality Tests:

  • Shapiro-Wilk: Best suited for small to medium-sized samples (typically n ≤ 2000; Stata may restrict it for larger samples).
  • Kolmogorov-Smirnov: Works for larger samples, but note that we’re using the sample mean and standard deviation to define the normal distribution here (this is the "Lilliefors" variant in practice).

内容的提问来源于stack exchange,提问作者Patudb

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.27 13:52:50