使用BY语句执行SAS多样本Logistic回归时遇有效观测缺失错误
Hey there! Let's work through that SAS error you're getting when using the BY statement to run logistic regressions for your 50 samples. That "no valid observations" error almost always means one or more of your BY groups doesn't have usable data to fit the model—here's how to diagnose and fix this step by step:
1. First, check for empty or incomplete BY groups
Start by confirming if any of your 50 samples have zero records, or if the grouping variable has missing values. Run this quick frequency check (replace sample_id with your actual BY group variable):
proc freq data=your_input_data; tables sample_id / missing; run;
This will show you exactly how many observations are in each group, including any groups marked as missing. If you see a group with zero counts, that's a clear culprit.
2. Validate your model variables per group
The error calls out missing values in the response, explanatory, frequency, or weight variables. Use proc means to spot groups where these variables are problematic:
proc means data=your_input_data n nmiss; by sample_id; var response_var pred_var1 pred_var2 pred_var3; /* List all variables used in your model */ run;
Keep an eye out for:
- Groups where the response variable is entirely missing
- Groups where a critical explanatory variable has no valid values
- If you're using a
freqorweightvariable: groups with non-positive values (0, negatives) or all missing entries
3. Filter out invalid groups before running regression
Once you've identified problematic groups, create a cleaned dataset that excludes them. Here's a way to do that with PROC SQL and a merge:
/* Step 1: Identify groups with valid data */ proc sql; create table valid_samples as select distinct sample_id from your_input_data where response_var is not missing and (pred_var1 is not missing or pred_var2 is not missing) /* Adjust to your predictors */ and (freq_var > 0 or freq_var is missing); /* Omit this line if not using freq/weight */ quit; /* Step 2: Keep only valid groups in your dataset */ data cleaned_data; merge your_input_data valid_samples; by sample_id; run; /* Step 3: Run logistic regression on cleaned data */ proc logistic data=cleaned_data; by sample_id; model response_var = pred_var1 pred_var2 pred_var3; run;
4. Watch out for small or unbalanced groups
Even if a group has some valid data, logistic regression can't run if the response variable has only one level (e.g., all 0s or all 1s) or if there's only 1 observation. Check this with another frequency table:
proc freq data=your_input_data; by sample_id; tables response_var; run;
Exclude any groups where the response has no variation—these can't produce a meaningful regression model.
Quick tip: Add the
NOBSoption to yourproc logisticstatement (likeproc logistic data=cleaned_data nobs;) to get a summary of how many groups were successfully processed, so you can confirm you're only working with valid samples.
内容的提问来源于stack exchange,提问作者Clément Hurel

