如何修改SAS代码生成正确复诊记录并对齐出入院时间?
问题描述
我有一个每行代表一次医院就诊的数据集,但当患者返回同一家医院时,系统会更新之前的出院时间(a5_disch_dttm)而非创建新记录。我需要生成额外的“复诊”行,使每次住院都有独立的入院(a1_first_dttm)和出院时间。
现有变量a5_tnf_fltu可标识转院目的地,能识别医院间转院,需要调整患者返回原医院时的日期时间:
- 当患者就诊路径为A→B→A时:
- 首次A院就诊:
a1_first_dttm保持不变,a5_disch_dttm改为下一次B院就诊的a1_first_dttm - B院就诊:保持不变
- 返回A院就诊:
a1_first_dttm改为上一次B院就诊的a5_disch_dttm,a5_disch_dttm使用原A院的出院时间
- 首次A院就诊:
现有数据(HAVE)
| rec_id | hospital | a1_first_dttm | a5_disch_dttm |
|---|---|---|---|
| 1 | A | 29AUG21:03:30 | 02SEP21:06:35 |
| 1 | B | 29AUG21:11:30 | 30AUG21:01:30 |
| 1 | A | . | 02SEP21:06:35 |
| 2 | A | 29JUN22:19:39 | 30JUN22:11:50 |
| 2 | B | 30JUN22:13:00 | 07JUL22:13:05 |
| 2 | C | 30JUN22:23:00 | 03JUL22:16:40 |
| 2 | B | . | 07JUL22:13:05 |
目标数据(WANT)
| rec_id | visit | hospital | a1_first_dttm | a5_disch_dttm |
|---|---|---|---|---|
| 1 | 1 | A | 29AUG21:03:30 | 29AUG21:11:30 |
| 1 | 2 | B | 29AUG21:11:30 | 30AUG21:01:30 |
| 1 | 3 | A | 30AUG21:01:30 | 02SEP21:06:35 |
| 2 | 1 | A | 29JUN22:19:39 | 30JUN22:11:50 |
| 2 | 2 | B | 30JUN22:13:00 | 30JUN22:23:00 |
| 2 | 3 | C | 30JUN22:23:00 | 03JUL22:16:40 |
| 2 | 4 | B | 03JUL22:16:40 | 07JUL22:13:05 |
当前SAS代码
data admit_out3; set admit_out2; * translate transfer disposition to transfer facility; if a5_tnf_fltu in (2,30) then trnsfr_fac = 1; else if a5_tnf_fltu in (4,31) then trnsfr_fac = 2; else if a5_tnf_fltu in (20,5,22,13,14,32) then trnsfr_fac = 3; else if a5_tnf_fltu in (21,3,23,15,16,33) then trnsfr_fac = 4; run; * Sort by patient and count; proc sort data=admit_out3; by rec_id count; run; * Generate return rows; data return_rows; set admit_out3; by rec_id; * Arrays to track visited facilities and their discharge datetimes; array fac_list[10] $50 _temporary_; array disch_list[10] _temporary_; retain fac_count; if first.rec_id then do; fac_count = 0; do i = 1 to dim(fac_list); fac_list[i] = ''; disch_list[i] = .; end; end; * Check if trnsfr_fac is already visited; found = 0; do i = 1 to fac_count; if trnsfr_fac = fac_list[i] then do; found = i; leave; end; end; if found > 0 then do; * Return visit; a1_sndfac = a1_recvfac; a1_recvfac = trnsfr_fac; Hospital = trnsfr_fac; a5_fl_dpo = .; a5_tnf_fltu = .; a5_disch_dttm = disch_list[found]; /* original discharge datetime */ a1_first_dttm = .; /* blank first_dttm for return */ count = count + 1; tot_visits = count; trnsfr_fac = .; output; end; * store the current facility for future reference; fac_count + 1; fac_list[fac_count] = a1_recvfac; disch_list[fac_count] = a5_disch_dttm; drop i found; run; * Combine original and return rows; data admit_out4; set admit_out3 return_rows(in=b); return_fac = b; run; * Sort final dataset; proc sort data=admit_out4; by rec_id count; run;
修改后的SAS代码及说明
核心思路
- 先按患者ID和就诊时间排序,确保就诊顺序正确
- 跟踪每个患者的就诊历史,识别需要拆分的重复医院记录
- 调整首次就诊的出院时间为下一次就诊的入院时间,同时生成复诊记录并填充正确的入院时间
* 第一步:整理基础数据,映射转院标识到医院代码,按患者+就诊时间排序; data admit_clean; set admit_out2; * 把转院标识映射为实际医院代码(根据你的业务规则调整); if a5_tnf_fltu in (2,30) then trnsfr_fac = 'A'; else if a5_tnf_fltu in (4,31) then trnsfr_fac = 'B'; else if a5_tnf_fltu in (20,5,22,13,14,32) then trnsfr_fac = 'C'; else if a5_tnf_fltu in (21,3,23,15,16,33) then trnsfr_fac = 'D'; run; proc sort data=admit_clean; by rec_id a1_first_dttm; run; * 第二步:生成完整就诊记录,调整出入院时间并创建复诊行; data want; set admit_clean; by rec_id; retain prev_disch_dttm visit_num; array orig_fac_disch[10] _temporary_; * 存储各医院的原始出院时间; array orig_fac[10] $50 _temporary_; retain fac_idx; if first.rec_id then do; visit_num = 0; prev_disch_dttm = .; fac_idx = 0; do i = 1 to dim(orig_fac); orig_fac[i] = ''; orig_fac_disch[i] = .; end; end; visit_num + 1; * 处理复诊记录(a1_first_dttm缺失的行); if missing(a1_first_dttm) then do; a1_first_dttm = prev_disch_dttm; * 匹配该医院的原始出院时间; do i = 1 to fac_idx; if orig_fac[i] = hospital then do; a5_disch_dttm = orig_fac_disch[i]; leave; end; end; end; * 处理首次就诊或普通新就诊; else do; * 检查当前医院是否有后续复诊记录; lead_hosp = lag(hospital, -1); lead_a1 = lag(a1_first_dttm, -1); if lead_hosp = hospital and missing(lead_a1) then do; * 有复诊,修改当前出院时间为下一次就诊的入院时间,保留原始出院时间; orig_disch = a5_disch_dttm; a5_disch_dttm = lead_a1; fac_idx + 1; orig_fac[fac_idx] = hospital; orig_fac_disch[fac_idx] = orig_disch; end; else do; * 无复诊,直接存储原始出院时间; fac_idx + 1; orig_fac[fac_idx] = hospital; orig_fac_disch[fac_idx] = a5_disch_dttm; end; end; * 更新上一次出院时间,用于下一条记录的入院时间填充; prev_disch_dttm = a5_disch_dttm; * 保留需要的变量并重命名; keep rec_id visit_num hospital a1_first_dttm a5_disch_dttm; rename visit_num=visit; drop i lead_hosp lead_a1 orig_disch fac_idx prev_disch_dttm; run; * 最终排序确保顺序正确; proc sort data=want; by rec_id visit; run;
关键修改点
- 时间排序:先按
rec_id和a1_first_dttm排序,保证就诊顺序的正确性,这是后续时间调整的基础 - 复诊识别:通过
missing(a1_first_dttm)自动识别返回原医院的复诊记录,填充入院时间为上一次就诊的出院时间 - 首次就诊时间调整:检查当前医院是否有后续复诊,若有则将首次就诊的出院时间改为下一次就诊的入院时间,同时用临时数组存储原始出院时间供复诊使用
- 临时数组存储:用临时数组记录各医院的原始出院时间,确保复诊记录能获取正确的最终出院时间
内容的提问来源于stack exchange,提问作者Sophia
相关产品推荐
相关产品推荐

