You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

能否将完整流程循环n次?SAS宏实现逻辑回归循环流程失败求助

Fixing Your SAS Macro for Repeating Logistic Regression Analysis

It looks like your macro has a few syntax issues and structural gaps that are keeping it from running properly. Let's break down the problems and walk through a complete, working solution tailored to your project needs.

First: Fixing Macro Syntax Errors

Your original macro had three critical syntax mistakes:

  1. Missing parameter definition: You tried passing 50 to %repeat() but didn't define a parameter in the macro header.
  2. Incorrect loop termination: You used end; instead of %end; to close the macro loop (%do requires %end for proper macro execution).
  3. Unstructured loop content: The code inside the loop lacked clear separation of data generation/analysis steps, leading to execution errors.

Working Macro Example

Here's a fully functional macro that generates 2 independent 1500-person samples per iteration, runs logistic regression on each, and saves results across 50 loops:

%macro repeat_logreg(num_loops=50);
  /* Create a master dataset to store all parameter estimates */
  proc datasets lib=work nolist;
    delete all_params;
  run;

  %do i = 1 %to &num_loops.;
    /* -------------------------- */
    /* Step 1: Generate Sample 1 */
    /* -------------------------- */
    data sample1_&i.;
      /* Set unique seed for reproducibility */
      call streaminit(&i.);
      do subject_id = 1 to 1500;
        /* Simulate predictor variables (adjust to match your project's covariates) */
        age = rand("NORMAL", 45, 10);
        treatment = rand("BINOMIAL", 0.5);
        /* Simulate logistic outcome */
        logit_prob = -2 + 0.03*age + 0.8*treatment;
        event_prob = exp(logit_prob)/(1 + exp(logit_prob));
        event = rand("BINOMIAL", event_prob);
        output;
      end;
    run;

    /* -------------------------- */
    /* Step 2: Generate Sample 2 */
    /* -------------------------- */
    data sample2_&i.;
      /* Different seed ensures independent sample */
      call streaminit(&i. + 1000);
      do subject_id = 1 to 1500;
        age = rand("NORMAL", 45, 10);
        treatment = rand("BINOMIAL", 0.5);
        logit_prob = -2 + 0.03*age + 0.8*treatment;
        event_prob = exp(logit_prob)/(1 + exp(logit_prob));
        event = rand("BINOMIAL", event_prob);
        output;
      end;
    run;

    /* -------------------------- */
    /* Step 3: Logistic Regression on Sample 1 */
    /* -------------------------- */
    proc logistic data=sample1_&i. noprint;
      model event(event='1') = age treatment;
      /* Save predicted probabilities */
      output out=preds1_&i. predicted=pred_event;
      /* Save parameter estimates to a dataset */
      ods output ParameterEstimates=params1_&i.;
    run;

    /* -------------------------- */
    /* Step 4: Logistic Regression on Sample 2 */
    /* -------------------------- */
    proc logistic data=sample2_&i. noprint;
      model event(event='1') = age treatment;
      output out=preds2_&i. predicted=pred_event;
      ods output ParameterEstimates=params2_&i.;
    run;

    /* -------------------------- */
    /* Step 5: Append Results to Master Dataset */
    /* -------------------------- */
    proc append base=all_params data=params1_&i. force;
      /* Add loop ID and sample identifier */
      sample_id = 1;
      loop_num = &i.;
    run;

    proc append base=all_params data=params2_&i. force;
      sample_id = 2;
      loop_num = &i.;
    run;

  %end;

  /* Optional: Clean up temporary sample datasets */
  proc datasets lib=work nolist;
    delete sample1_: sample2_: preds1_: preds2_: params1_: params2_:;
  run;
%mend repeat_logreg;

/* Execute the macro with 50 loops */
%repeat_logreg(num_loops=50);

Key Features & Tips

  • Reproducibility: Using call streaminit(&i.) sets a unique seed per loop, so you can replicate results later. Adding an offset (+1000) for the second sample ensures the two samples are independent.
  • Result Storage: The ods output ParameterEstimates captures regression coefficients for each sample, and proc append combines all results into a single dataset (all_params) for easy post-loop analysis.
  • Clean Code: Using noprint in proc logistic suppresses redundant output, and the final proc datasets cleans up temporary files to save workspace space.
  • Flexibility: Adjust the data generation steps to match your actual covariates and outcome (replace the simulated variables with your sampling logic if you're using an existing population dataset).

If You're Sampling from an Existing Dataset

If you're not simulating data but drawing samples from a pre-existing dataset (e.g., my_population), replace the data steps with proc surveyselect:

/* Replace sample generation data steps with this */
proc surveyselect data=my_population out=sample1_&i. method=srs n=1500 seed=&i.;
run;

proc surveyselect data=my_population out=sample2_&i. method=srs n=1500 seed=&i.+1000;
run;

Content of this question comes from Stack Exchange, asked by Clément Hurel

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 09:42:10