You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Python中创建可复用函数实现指定列获取与查询条件动态应用

Pandas 可复用列选取与条件过滤函数实现

原有代码中过滤条件的写法存在逻辑错误:["status = 'closed'" and "customer_type = 'R'"] 会先执行两个字符串的布尔与运算,最终列表中只会保留第二个字符串,无法实现多条件同时生效的需求。

函数实现

import pandas as pd

def select_filter_df(df, columns_list=None, filter_conditions=None):
    """
    对输入DataFrame执行指定列选取、条件过滤后返回结果
    参数说明:
        df: 输入的Pandas DataFrame对象
        columns_list: 要选取的列名列表,传None则默认返回所有列
        filter_conditions: 过滤条件字符串列表,多个条件默认用and拼接,传None则默认不执行过滤
    返回值:
        处理完成的DataFrame对象
    """
    result_df = df.copy()
    # 列选取逻辑
    if columns_list is not None:
        # 校验合法列名,避免传入不存在的列报错
        valid_cols = [col for col in columns_list if col in result_df.columns]
        if not valid_cols:
            raise ValueError("所有传入的列名均不存在于当前DataFrame中")
        result_df = result_df[valid_cols]
    # 条件过滤逻辑
    if filter_conditions and len(filter_conditions) > 0:
        # 自动拼接多个过滤条件
        query_expression = " and ".join(filter_conditions)
        result_df = result_df.query(query_expression)
    return result_df

使用示例

# 读取数据
df = read_df() # 替换为你自身的读数据逻辑

# 定义入参
columns_list = ["ticket_start_time", "ticket_end_time", "status", "customer_type", "ticket_type"]
filter_conditions =  ["status == 'Por Acción'", "customer_type == 'R'"]

# 调用函数获取结果
processed_df = select_filter_df(df, columns_list, filter_conditions)

优化点说明

  • 内置列名校验逻辑,避免传入不存在的列直接触发报错
  • 支持不传列列表/过滤条件的场景,灵活适配不同使用需求
  • 多条件自动拼接,无需硬编码写逻辑,新增条件仅需往filter_conditions列表中添加对应字符串即可
  • 先做列选取再做过滤,数据量大时可减少过滤计算量,提升运行性能

内容的提问来源于stack exchange,提问作者Pavithra Kannan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.24 09:27:02