Python医疗数据处理:同日双病症排查及年度日期排除方案求助
医疗数据集清洗解决方案
假设你的数据集已用pandas加载为DataFrame,核心列包含:
PersonID: 用户唯一标识Date: 记录日期(建议先转为datetime类型)HealthCondition: 健康状态(取值为'Health Condition 1'、'Health Condition 2'或其他)
以下是分步解决代码:
步骤1:标记并排除同一人同一天同时存在两种病症的条目
import pandas as pd # 转换日期列格式,方便后续年份提取 df['Date'] = pd.to_datetime(df['Date']) # 分组检查每个用户每日是否同时有两种病症 has_both_conditions = df.groupby(['PersonID', 'Date'])['HealthCondition'].apply( lambda x: ('Health Condition 1' in x.values) and ('Health Condition 2' in x.values) ).reset_index(name='has_both') # 合并标记结果到原数据集 df = df.merge(has_both_conditions, on=['PersonID', 'Date'], how='left') # 提取第一类需要排除的条目索引 exclude_type1 = df[df['has_both'] == True].index
步骤2:标记并排除目标人群对应年份的所有条目
# 从第一类排除条目中,提取患有Health Condition 2的记录,获取对应年份 hc2_target_records = df[(df['has_both'] == True) & (df['HealthCondition'] == 'Health Condition 2')] # 整理每个用户需要排除的年份(一个用户可能有多个年份) user_exclude_years = hc2_target_records.groupby('PersonID')['Date'].apply( lambda x: x.dt.year.unique() ).reset_index(name='exclude_years') # 展开年份列表,方便后续匹配 user_exclude_years = user_exclude_years.explode('exclude_years') # 合并年份信息到原数据集,标记记录所属年份 df = df.merge(user_exclude_years, on='PersonID', how='left') df['record_year'] = df['Date'].dt.year # 提取第二类需要排除的条目索引 exclude_type2 = df[(df['exclude_years'] == df['record_year'])].index
步骤3:生成最终清洗后的数据集
# 合并两类排除条目索引 all_exclude_indices = exclude_type1.union(exclude_type2) # 删除排除条目并清理临时列 cleaned_df = df.drop(all_exclude_indices).drop(columns=['has_both', 'exclude_years', 'record_year'])
注意事项
- 如果
HealthCondition的实际取值与示例不同,替换代码中对应的字符串即可 - 若日期列格式特殊,需在
pd.to_datetime中指定format参数,例如pd.to_datetime(df['Date'], format='%Y/%m/%d')
内容的提问来源于stack exchange,提问作者user20743916
相关产品推荐
相关产品推荐

