You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

求助:基于已有疾病症状DataFrame构造新DataFrame的方法

实现DataFrame格式转换的可行方案

步骤1:补全缺失字段值

原始DataFrame中desease和occurences列仅每组症状的首行有值,其余为空,用前向填充补全这些缺失值:

import pandas as pd

# 假设原始数据已读入为df
df['desease'] = df['desease'].ffill()
df['occurences'] = df['occurences'].ffill()

步骤2:对症状列生成独热编码

利用pd.get_dummies将症状列转换为二进制标记列,每个症状对应一列,出现标记为1,未出现为0:

symptom_dummies = pd.get_dummies(df['symptoms'], prefix='', prefix_sep='')

步骤3:按疾病分组聚合

以desease和occurences为分组依据,对独热编码后的列取最大值(确保同一疾病的所有症状都标记为1):

aggregated = symptom_dummies.groupby([df['desease'], df['occurences']]).max().reset_index()

步骤4:调整列顺序

将occurences和desease列移至末尾,与目标格式对齐:

# 提取所有症状列
symptom_cols = [col for col in aggregated.columns if col not in ['desease', 'occurences']]
# 重新排列列
final_df = aggregated[symptom_cols + ['occurences', 'desease']]

运行上述代码后,即可得到你需要的目标DataFrame格式。

内容的提问来源于stack exchange,提问作者henri azemena

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.12 23:52:50