You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何高效按对象属性筛选对象数组并优化Python患者数据处理代码

优化方案:提升患者治疗计划处理效率

问题根源

你的代码效率低主要来自两个核心问题:

  • 用列表存储患者,每次查找需遍历整个列表(any(x.pat_id == index[4] for x in patients)),时间复杂度O(n),大数据量下重复遍历会严重拖慢速度。
  • 患者的plans用列表存储,判断计划是否存在需遍历列表(index[11] not in pat.plans),O(k)的时间复杂度在重复操作多时会导致效率骤降。

优化代码

1. 修改Patient类,用字典+集合优化存储

class Patient:
    # 用字典替代列表,通过pat_id直接映射到患者对象,实现O(1)查找
    all_patients = {}

    def __init__(self, pat_id):
        self.pat_id = pat_id
        # 用集合存储治疗计划,自动去重,成员判断/添加操作都是O(1)
        self.plans = set()
        Patient.all_patients[pat_id] = self

2. 日志处理逻辑简化优化

total_plans = 0

# 处理单条日志数据(index为当前行解析后的数据)
pat_id = index[4]
plan_id = index[11]

# 直接通过字典获取患者,不存在则自动创建
current_patient = Patient.all_patients.get(pat_id)
if not current_patient:
    current_patient = Patient(pat_id)

# 集合自动忽略重复计划,只需判断是否新增成功
if plan_id not in current_patient.plans:
    current_patient.plans.add(plan_id)
    total_plans += 1

# 输出结果
print(len(Patient.all_patients))
print(total_plans)

优化点说明

  • 患者字典映射:把all_patients从列表改成字典,通过患者ID直接定位对象,彻底避免遍历列表的开销,查找速度从O(n)降到O(1)。
  • 集合存储计划:集合的成员判断和添加操作都是常数时间,自动去重,不用手动写重复检查逻辑,既简化代码又提升效率。
  • 减少遍历次数:原来的代码要先遍历列表判断患者是否存在,再遍历列表找患者添加计划,现在一次字典查找就能完成患者定位,减少了冗余操作。

兼容原有类的替代方案(如果无法修改Patient类)

如果因为其他依赖不能修改原类,可在外部维护患者字典,同时保留原有列表存储计划:

# 外部维护患者ID到对象的映射
patient_map = {}
total_plans = 0

# 处理单条日志
pat_id = index[4]
plan_id = index[11]

# 获取或创建患者
current_patient = patient_map.get(pat_id)
if not current_patient:
    current_patient = patient(pat_id)
    patient_map[pat_id] = current_patient

# 判断计划是否重复并添加
if plan_id not in current_patient.plans:
    current_patient.plans.append(plan_id)
    total_plans += 1

内容的提问来源于stack exchange,提问作者Thomas

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.27 19:33:11