如何匹配SAS与Python的随机种子以获得一致抽样结果
Great question! Yes, it's absolutely possible to get identical sampling results between SAS and Python when using the same seed—you just need to align how the two tools handle random number generation (RNG), since their default implementations don't match out of the box. Let's break this down step by step.
Step 1: Confirm SAS's Random Number Generator
First off, check which RNG your SAS code is using, because newer SAS versions use a different generator than older ones. Run this quick SAS snippet to find out:
proc options option=randomnumbergenerator; run;
Most modern SAS (9.4+) uses MERSENNE_TWISTER, but if you're on an older version, it might be RANUNI (a linear congruential generator). We'll cover both scenarios below.
Step 2: Replicate Mersenne Twister (SAS 9.4+) in Python
SAS's Mersenne Twister implementation is compatible with the standard one used in Python's numpy library, but there's a small catch with how seeds are initialized. Here's how to mirror your SAS sampling:
Example SAS Sampling Code
Suppose your SAS code looks like this (using the modern stream-based RNG):
data sampled; set your_dataset; call streaminit(12345); /* Set seed */ if rand('uniform') <= 0.3 then output; /* Keep 30% of rows */ run;
Equivalent Python Code
Use numpy's MT19937 generator, initialized with the same seed as SAS. This will produce matching uniform random numbers:
import numpy as np import pandas as pd # Load your dataset (make sure row order matches SAS exactly!) df = pd.read_csv("your_dataset.csv") # Initialize numpy's RNG to match SAS's Mersenne Twister rng = np.random.MT19937(seed=12345) random_state = np.random.RandomState(rng) # Generate uniform random numbers identical to SAS df['rand_num'] = random_state.uniform(0, 1, size=len(df)) # Filter to get the same 30% sample sampled_df = df[df['rand_num'] <= 0.3].reset_index(drop=True)
Pro tip: Test this with a tiny dataset first (e.g., 5 rows) to confirm the random numbers match between SAS and Python before scaling up.
Step 3: Replicate RANUNI (Older SAS Versions)
If your SAS code uses the older ranuni function instead of streaminit, you'll need to manually replicate the linear congruential generator (LCG) that ranuni uses. SAS's ranuni follows this formula:seed = mod(seed * 397204094 + 1, 2**31)
The resulting random number is seed / 2**31.
Example SAS Code with RANUNI
data sampled; set your_dataset; retain seed 12345; /* Initialize seed */ seed = ranuni(seed); /* Generate next random number */ if seed <= 0.3 then output; run;
Equivalent Python Code
Implement the LCG directly to match ranuni's behavior:
import pandas as pd df = pd.read_csv("your_dataset.csv") # Initialize seed to match SAS seed = 12345 rand_nums = [] # Generate random numbers using SAS's RANUNI formula for _ in range(len(df)): seed = (seed * 397204094 + 1) % (2**31) rand_nums.append(seed / (2**31)) df['rand_num'] = rand_nums sampled_df = df[df['rand_num'] <= 0.3].reset_index(drop=True)
This will produce exactly the same sequence of random numbers as SAS's ranuni with seed 12345.
Critical Things to Remember
- Row Order is Everything: Make sure the order of rows in your Python dataset is identical to what SAS is using. If you're reading from a file, double-check that pandas isn't reordering rows (it won't by default, but it's worth confirming).
- Match Distributions: If you're using non-uniform random functions (like
rand('normal')in SAS), you'll need to replicate those distributions using the same RNG, since SAS and Python might generate normals or other distributions differently even with the same seed. - Test Small First: Always validate with a tiny dataset—generate 5-10 random numbers in both tools to confirm they match before working with your full dataset.
内容的提问来源于stack exchange,提问作者Franco Garrido

