SAS-多变量卡方独立性检验代码需求:二分类变量关联分析
SAS Code for Batch Chi-Square Tests & Sorted Correlation Results
Got it, let's work through this SAS code request. You need to run chi-square independence tests between your binary variable v1 and over 200 categorical variables (v2 to v200+), exclude the ID column, then sort the results by chi-square statistic to find the variables most correlated with v1. Here's a solid, scalable solution:
Full SAS Code
/* -------------------------- Step 1: Identify target variables -------------------------- */ /* Get list of all variables in your dataset */ proc contents data=your_data out=var_list noprint; run; /* Filter to keep only categorical variables (exclude ID and v1) */ /* Adjust TYPE=2 to TYPE=1 if your categorical variables are numeric */ data var_list; set var_list; where upcase(name) not in ('ID', 'V1') and type=2; keep name; run; /* Convert variable list to a macro variable for looping */ proc sql noprint; select name into :cat_vars separated by ' ' from var_list; quit; /* -------------------------- Step 2: Prepare results dataset -------------------------- */ /* Create empty dataset to store chi-square test outputs */ data chi_results; length var_name $32 chi_square p_value df 8; stop; /* Stop immediately to create structure without rows */ run; /* -------------------------- Step 3: Macro to run batch chi-square tests -------------------------- */ %macro run_chi_tests; %let i = 1; %let current_var = %scan(&cat_vars., &i.); /* Loop through each categorical variable */ %do %while(¤t_var. ne); /* Run chi-square test between v1 and current variable */ proc freq data=your_data noprint; tables v1 * ¤t_var. / chisq; output out=temp_results chisq; /* Export test stats to temp dataset */ run; /* Clean up temp results and add variable name */ data temp_results; set temp_results; var_name = "¤t_var."; keep var_name chi_square p_value df; run; /* Append temp results to master results dataset */ proc append base=chi_results data=temp_results; run; /* Move to next variable in the list */ %let i = %eval(&i. + 1); %let current_var = %scan(&cat_vars., &i.); %end; %mend run_chi_tests; /* Execute the macro */ %run_chi_tests; /* -------------------------- Step 4: Sort and display results -------------------------- */ /* Sort results by chi-square statistic (descending = most correlated first) */ proc sort data=chi_results descending; by chi_square; run; /* Print the sorted results with clear labels */ proc print data=chi_results label noobs; label var_name = 'Categorical Variable' chi_square = 'Chi-Square Statistic' p_value = 'P-Value' df = 'Degrees of Freedom'; title 'Top Variables Correlated with v1 (Sorted by Chi-Square)'; run;
Key Notes & Adjustments
- Variable Type: The code assumes your categorical variables are character-type (
type=2inproc contents). If your categorical variables are numeric (e.g., coded 1/2/3), changetype=2totype=1in thevar_listdata step. - ID Column: The code excludes
ID(case-insensitive viaupcase(name)). If your ID column has a different name, update thenot inlist. - Small Cell Counts: If any variable has expected cell counts <5, the chi-square test may be unreliable. You can add a check for this, or modify the
tablesstatement to includeexact fisherif you need Fisher's exact test instead for small samples. - Dataset Name: Replace
your_datawith the actual name of your input dataset.
How It Works
- Variable Identification: We first extract all relevant categorical variables (excluding ID and v1) to avoid manual listing of 200+ variables.
- Batch Testing: A macro loops through each variable, runs the chi-square test, and stores the key statistics (chi-square value, p-value, degrees of freedom) along with the variable name.
- Sorted Results: Finally, we sort the results by chi-square statistic in descending order so the variables most correlated with
v1appear at the top.
内容的提问来源于stack exchange,提问作者bartek09
相关产品推荐
相关产品推荐

