You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python3.7/Pandas中如何引用全局配置文件实现DataFrame数据过滤

实现思路
  • 首先解析配置规则:区分为等值匹配和运算符匹配两种模式,配置值为非列表类型时默认执行等值匹配,为列表时第一个元素为运算符、第二个为匹配阈值
  • 逐行应用过滤条件:遍历配置字典的每一条规则,依次对DataFrame做条件筛选,每一步保留符合规则的行进入下一轮过滤
  • 生成过滤流水记录:每一步过滤都记录字段名、过滤规则、过滤前后行数、剔除行数,最终生成waterfall统计表
  • 返回指定结果:最终输出过滤后的合格数据集、过滤流水表、主键级别的匹配计数
代码实现

首先确认已安装pandas依赖,将以下代码放入src/utils/utilities.py文件中即可:

import pandas as pd

def filter_me(raw_df: pd.DataFrame, filter_config: dict, primary_key: str) -> tuple[pd.DataFrame, pd.DataFrame, dict]:
    # 初始化中间变量
    current_df = raw_df.copy()
    total_raw_rows = len(current_df)
    waterfall_records = []
    # 运算符映射表,可根据需求扩展更多操作符
    op_map = {
        '>': 'gt',
        '>=': 'ge',
        '<': 'lt',
        '<=': 'le',
        '!=': 'ne',
        '==': 'eq'
    }

    # 遍历所有过滤规则
    for col, rule in filter_config.items():
        # 字段存在性校验
        if col not in current_df.columns:
            raise ValueError(f"过滤字段{col}不存在于输入DataFrame中")
        pre_filter_count = len(current_df)

        # 区分规则类型做匹配
        if isinstance(rule, list):
            op, threshold = rule
            if op not in op_map:
                raise ValueError(f"不支持的运算符{op},当前仅支持{list(op_map.keys())}")
            # 调用pandas内置比较方法生成过滤掩码
            filter_mask = getattr(current_df[col], op_map[op])(threshold)
        else:
            # 等值匹配逻辑
            filter_mask = current_df[col] == rule

        # 应用过滤规则
        current_df = current_df[filter_mask].reset_index(drop=True)
        post_filter_count = len(current_df)

        # 记录当前步骤的过滤流水
        waterfall_records.append({
            "过滤字段": col,
            "过滤规则": f"{op} {threshold}" if isinstance(rule, list) else f"== {rule}",
            "过滤前行数": pre_filter_count,
            "过滤后行数": post_filter_count,
            "剔除行数": pre_filter_count - post_filter_count,
            "剔除占比(%)": round((pre_filter_count - post_filter_count)/pre_filter_count*100, 2) if pre_filter_count else 0
        })

    # 生成返回结果
    df_waterfall = pd.DataFrame(waterfall_records)
    # mc为统计指标字典,可根据实际业务需求调整字段
    mc = {
        "原始总行数": total_raw_rows,
        "过滤后总行数": len(current_df),
        "原始主键去重数": raw_df[primary_key].nunique(),
        "过滤后主键去重数": current_df[primary_key].nunique(),
        "总保留占比(%)": round(len(current_df)/total_raw_rows*100, 2) if total_raw_rows else 0
    }
    return current_df, df_waterfall, mc
兼容说明

以上实现完全兼容你现有的调用逻辑,传入示例中的global_config_eam配置即可自动完成所有过滤规则的匹配。如果需要扩展in、between等更复杂的过滤逻辑,只需修改op_map映射表和对应规则解析逻辑即可。

内容的提问来源于stack exchange,提问作者rageousquitter

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.27 01:45:04