如何读取含变量名=值格式的SAS输入数据并生成目标数据集?
Got it, let's break down how to read this variable=value formatted data in SAS to get exactly the dataset you want (with just name and math columns).
First, your input data is structured as key-value pairs where each person's attributes are listed in a single line. The tricky part is handling names with spaces (like "Mary Louise") and only extracting the variables you care about.
Method 1: Extract Values Directly from the Full Line
This is the simplest approach if you only need name and math—we'll read the entire line as text, then pull out the values we need using string manipulation functions:
data want; /* Define variable lengths upfront to avoid truncating names */ length name $20 math 8; infile datalines dlm=' ' missover; /* Hold the entire line in the input buffer so we can access it with _infile_ */ input @; /* Extract name: Replace "name=" with a delimiter, then grab the first segment (including spaces) */ name = scan(tranwrd(_infile_, 'name=', '|'), 1, ' |'); /* Extract math: Replace "math=" with a delimiter, convert the string to a number */ math = input(scan(tranwrd(_infile_, 'math=', '|'), 1, ' |'), 8.); /* Keep only the variables we need */ keep name math; /* Your raw data */ datalines; name=Linda english=95 math=94 science=90 name=Susan math=88 english=91 science=90 name=Mary Louise math=90 english=84 science=81 ; run; /* Verify the output */ proc print data=want; run;
How this works:
input @;keeps the entire line in the input buffer, so we can use the special_infile_variable to read the full text of the line.tranwrd(_infile_, 'name=', '|')replaces the "name=" prefix with a pipe (|), making it easy to split the line with thescanfunction. We split on both|and spaces to capture the full name (even with spaces).- For
math, we do the same but convert the extracted string to a numeric value usinginput().
Method 2: Generalized Key-Value Parsing (Great for Future Expansion)
If you might need to extract other variables later (like english or science), this loop-based method is more flexible. It reads each key-value pair one by one and assigns values to variables:
data want; length key $10 value $20 name $20 math 8; infile datalines dlm=' ' dsd missover; /* Loop through all key-value pairs in the line */ do until (eof); input key value @; /* Assign values to variables based on the key */ if key = 'name=' then name = value; else if key = 'math=' then math = input(value, 8.); /* Handle names with spaces: if we haven't finished the name, append the next segment */ else if missing(name) then name = catx(' ', name, key); /* Check if we've reached the last pair in the line */ eof = (scan(_infile_, -1, ' ') = key); end; keep name math; datalines; name=Linda english=95 math=94 science=90 name=Susan math=88 english=91 science=90 name=Mary Louise math=90 english=84 science=81 ; run;
How this works:
- We loop through each key-value pair (split by spaces). When we hit
name=, we start storing the name. If the next "key" isn't a known variable (like "Louise"), we append it to the name since it's part of the full name. - When we hit
math=, we convert the value to a number and assign it to themathvariable.
Expected Output
Either method will produce exactly the dataset you wanted:
| Obs | name | math |
|---|---|---|
| 1 | Linda | 94 |
| 2 | Susan | 88 |
| 3 | Mary Louise | 90 |
内容的提问来源于stack exchange,提问作者Sujit Patnaik

