You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Pandas绘制DataFrame中patient_id列的分布可视化求助

实现Patient ID列的分布可视化方案

别担心,这个需求其实很常见,我来一步步教你搞定!首先我们要明确:你要的是每个患者对应的记录数量的分布——也就是统计有多少患者有1条记录、多少有2条,以此类推,再把这个统计结果可视化出来。

第一步:先统计每个患者的记录数

首先用Pandas对patient_id列做计数统计,这是可视化的基础:

import pandas as pd

# 假设你的DataFrame名为df
patient_counts = df['patient_id'].value_counts()
# patient_counts的结构:索引是patient_id,对应的值是该患者的记录条数

第二步:选择合适的可视化方法

根据你的数据集大小和需求,这里推荐3种常用的方式:

1. 直方图(最适合看整体分布趋势)

如果你的患者数量很多,直方图能清晰展示「不同记录数区间内的患者数量」,还可以叠加密度曲线辅助观察:

import seaborn as sns
import matplotlib.pyplot as plt

sns.histplot(patient_counts, bins=15, kde=True)
plt.xlabel('Number of Records per Patient')
plt.ylabel('Number of Patients')
plt.title('Distribution of Patient Record Counts')
plt.show()
  • bins参数可以调整区间数量,根据你的数据分布灵活调整即可
  • kde=True会添加一条密度曲线,帮你更直观地捕捉分布的峰值

2. 条形图(适合患者数量较少的场景)

如果你的患者总数不多,条形图可以直接展示每个患者的具体记录数:

patient_counts.plot(kind='bar', figsize=(12,6))
plt.xlabel('Patient ID')
plt.ylabel('Number of Records')
plt.title('Number of Records per Patient')
plt.xticks(rotation=45)  # 旋转x轴标签避免重叠
plt.show()
  • figsize可以调整图的大小,适配你的显示需求

3. 箱线图(快速看统计特征)

如果你想快速了解记录数的中位数、四分位数以及异常值情况,箱线图是个不错的选择:

sns.boxplot(x=patient_counts)
plt.xlabel('Number of Records per Patient')
plt.title('Box Plot of Patient Record Counts')
plt.show()

额外小技巧

如果你的记录数差异极大(比如有的患者有几百条记录,有的只有1条),可以给x轴加上对数刻度,让分布更清晰:

sns.histplot(patient_counts, bins=15, kde=True)
plt.xscale('log')
plt.xlabel('Number of Records per Patient (Log Scale)')
plt.ylabel('Number of Patients')
plt.title('Distribution of Patient Record Counts (Log Scale)')
plt.show()

内容的提问来源于stack exchange,提问作者Mike S

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 09:49:34