如何基于DataFrame实现灵活易修改的多条件样本计数
多条件灵活统计DataFrame样本数的实用方案
针对你需要频繁调整多条件(含不等判断)统计样本数的需求,推荐用条件字典+通用筛选函数的方案,既能快速修改规则,又能高效处理复杂逻辑。
核心方案:用字典存储条件,动态生成筛选逻辑
把每个列的判断规则存入字典,修改条件只需要更新字典内容,不用改动核心筛选代码。
1. 定义条件字典
以你给出的示例条件(Genepanel == PID、Causative != Negative、Gender == Male且Fenotype == Hypogamma)为例,字典可以这么写:
conditions = { "Genepanel": lambda x: x == "PID", "Causative": lambda x: x != "Negative", "Gender": lambda x: x == "Male", "Fenotype": lambda x: x == "Hypogamma" }
lambda表达式能灵活支持各种判断:比如范围判断lambda x: x > 18、多值匹配lambda x: x in ["SCID", "Hypogamma"]、包含判断lambda x: "突变" in x等。
2. 编写通用统计函数
写一个复用性强的函数,自动遍历条件字典生成筛选掩码,最后返回符合条件的样本数:
import pandas as pd def count_matching_samples(df, conditions): # 初始化全为True的掩码 mask = pd.Series([True] * len(df)) # 遍历每个条件,逐步缩小筛选范围 for col, rule in conditions.items(): mask &= df[col].apply(rule) # 返回符合条件的行数总和 return mask.sum()
3. 调用函数统计结果
假设你的患者数据DataFrame名为patient_df,直接传入参数即可得到结果:
result_count = count_matching_samples(patient_df, conditions) print(result_count) # 示例数据中输出1
优化:针对大型DataFrame的向量化版本
如果你的数据量极大(数万行以上),可以用向量化操作替代apply提升速度,条件字典的格式稍作调整:
# 用元组表示不等/范围类条件,字符串表示等于条件 conditions_optimized = { "Genepanel": "PID", "Causative": ("!=", "Negative"), "Gender": "Male", "Fenotype": "Hypogamma" } def count_matching_samples_fast(df, conditions): mask = pd.Series([True] * len(df)) for col, rule in conditions.items(): if isinstance(rule, str): mask &= (df[col] == rule) elif isinstance(rule, tuple): op, value = rule if op == "!=": mask &= (df[col] != value) elif op == ">": mask &= (df[col] > value) # 可扩展支持<、>=、<=等操作符 return mask.sum()
扩展:处理跨列关联条件
如果需要用到列之间的逻辑(比如Age > 18 且 Causative != Negative),可以单独生成跨列掩码后合并:
# 基础列条件 basic_conditions = {"Gender": "Male", "Genepanel": "PID"} # 跨列额外条件 cross_mask = (patient_df["Age"] > 18) & (patient_df["Causative"] != "Negative") # 合并掩码并统计 total_mask = count_matching_samples_fast(patient_df, basic_conditions) & cross_mask final_count = total_mask.sum()
这个方案的优势在于:
- 易维护:新增/修改条件只需调整字典,不用重构筛选逻辑
- 灵活性强:支持几乎所有常见的判断规则
- 性能适配:可根据数据量选择普通版或向量化版
内容的提问来源于stack exchange,提问作者Sofie
相关产品推荐
相关产品推荐

