基于前3行其他列值创建Y/N型标记列
Alright, let's break down how to recreate those two flag columns—for diabetes and hypertension—from your mortality dataset, focusing only on 2010-2016 data and checking the first 3 ENICON columns (ENICON1 to ENICON3) as specified.
Step 1: Filter to Target Years
First, narrow down your dataset to only the years you care about:
import pandas as pd # Replace 'mortality_data' with your actual DataFrame name filtered_data = mortality_data[(mortality_data['year'] >= 2010) & (mortality_data['year'] <= 2016)]
Step 2: Define Your Target Codes
You'll need to map the ICD codes (or whatever encoding your dataset uses) for diabetes and hypertension. I'll use common ICD-10 codes as examples—swap these out for the actual values in your ENICON columns:
# Example ICD-10 codes for diabetes diabetes_codes = ['E10', 'E11', 'E12', 'E13', 'E14'] # Example ICD-10 codes for hypertension hypertension_codes = ['I10', 'I11', 'I12', 'I13']
Step 3: Create the Flag Columns
The core logic here is: check if any of the first 3 ENICON columns contain a target code, then mark the flag as 1 (yes) or 0 (no).
Diabetes Flag
# Check ENICON1-3 for diabetes codes; return 1 if any match, 0 otherwise filtered_data['diabetes_flag'] = filtered_data[['ENICON1', 'ENICON2', 'ENICON3']].isin(diabetes_codes).any(axis=1).astype(int)
Hypertension Flag
# Repeat the logic for hypertension filtered_data['hypertension_flag'] = filtered_data[['ENICON1', 'ENICON2', 'ENICON3']].isin(hypertension_codes).any(axis=1).astype(int)
Step 4: Verify Your Results
To make sure the flags are working as expected, spot-check some rows:
# View the first 10 rows with the relevant columns print(filtered_data[['ENICON1', 'ENICON2', 'ENICON3', 'diabetes_flag', 'hypertension_flag']].head(10))
SAS Logic Reference
For context, here's what the equivalent SAS code would look like, matching your original logic:
data filtered_data; set mortality_data; where year between 2010 and 2016; /* Initialize flags to 0 (no) */ diabetes_flag = 0; hypertension_flag = 0; /* Set diabetes flag if any of first 3 ENICON columns have a match */ if ENICON1 in ('E10','E11','E12','E13','E14') or ENICON2 in ('E10','E11','E12','E13','E14') or ENICON3 in ('E10','E11','E12','E13','E14') then diabetes_flag = 1; /* Set hypertension flag similarly */ if ENICON1 in ('I10','I11','I12','I13') or ENICON2 in ('I10','I11','I12','I13') or ENICON3 in ('I10','I11','I12','I13') then hypertension_flag = 1; run;
内容的提问来源于stack exchange,提问作者Mariana Cendon

