You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于DataFrame实现灵活易修改的多条件样本计数

多条件灵活统计DataFrame样本数的实用方案

针对你需要频繁调整多条件(含不等判断)统计样本数的需求,推荐用条件字典+通用筛选函数的方案,既能快速修改规则,又能高效处理复杂逻辑。

核心方案:用字典存储条件,动态生成筛选逻辑

把每个列的判断规则存入字典,修改条件只需要更新字典内容,不用改动核心筛选代码。

1. 定义条件字典

以你给出的示例条件(Genepanel == PID、Causative != Negative、Gender == Male且Fenotype == Hypogamma)为例,字典可以这么写:

conditions = {
    "Genepanel": lambda x: x == "PID",
    "Causative": lambda x: x != "Negative",
    "Gender": lambda x: x == "Male",
    "Fenotype": lambda x: x == "Hypogamma"
}

lambda表达式能灵活支持各种判断:比如范围判断lambda x: x > 18、多值匹配lambda x: x in ["SCID", "Hypogamma"]、包含判断lambda x: "突变" in x等。

2. 编写通用统计函数

写一个复用性强的函数,自动遍历条件字典生成筛选掩码,最后返回符合条件的样本数:

import pandas as pd

def count_matching_samples(df, conditions):
    # 初始化全为True的掩码
    mask = pd.Series([True] * len(df))
    # 遍历每个条件,逐步缩小筛选范围
    for col, rule in conditions.items():
        mask &= df[col].apply(rule)
    # 返回符合条件的行数总和
    return mask.sum()

3. 调用函数统计结果

假设你的患者数据DataFrame名为patient_df,直接传入参数即可得到结果:

result_count = count_matching_samples(patient_df, conditions)
print(result_count)  # 示例数据中输出1

优化:针对大型DataFrame的向量化版本

如果你的数据量极大(数万行以上),可以用向量化操作替代apply提升速度,条件字典的格式稍作调整:

# 用元组表示不等/范围类条件,字符串表示等于条件
conditions_optimized = {
    "Genepanel": "PID",
    "Causative": ("!=", "Negative"),
    "Gender": "Male",
    "Fenotype": "Hypogamma"
}

def count_matching_samples_fast(df, conditions):
    mask = pd.Series([True] * len(df))
    for col, rule in conditions.items():
        if isinstance(rule, str):
            mask &= (df[col] == rule)
        elif isinstance(rule, tuple):
            op, value = rule
            if op == "!=":
                mask &= (df[col] != value)
            elif op == ">":
                mask &= (df[col] > value)
            # 可扩展支持<、>=、<=等操作符
    return mask.sum()

扩展:处理跨列关联条件

如果需要用到列之间的逻辑(比如Age > 18 且 Causative != Negative),可以单独生成跨列掩码后合并:

# 基础列条件
basic_conditions = {"Gender": "Male", "Genepanel": "PID"}
# 跨列额外条件
cross_mask = (patient_df["Age"] > 18) & (patient_df["Causative"] != "Negative")
# 合并掩码并统计
total_mask = count_matching_samples_fast(patient_df, basic_conditions) & cross_mask
final_count = total_mask.sum()

这个方案的优势在于:

  • 易维护:新增/修改条件只需调整字典,不用重构筛选逻辑
  • 灵活性强:支持几乎所有常见的判断规则
  • 性能适配:可根据数据量选择普通版或向量化版

内容的提问来源于stack exchange,提问作者Sofie

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.19 15:03:13