You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于另一列子串填充列并高效管理80+条件?

用配置文件管理条件与结果的解决方案

当然可以通过配置文件来管理这些条件和对应的结果值,这能彻底解决80+条件维护繁琐的问题,新增或修改条件只需编辑配置文件,无需改动核心代码。下面给出两种实用的实现方案:

方案1:使用CSV配置文件(适合简单条件)

1. 编写配置文件

创建config.csv,每行对应一个条件-结果对,格式如下(支持单条件或多条件组合):

Set,==,Z,Type,==,A,yellow
Set,==,Z,Type,==,B,blue
Set,==,X,Type,==,B,purple
Type,==,C,,default,,black
  • 前六列用于定义条件:col1, op1, val1对应第一个列的匹配规则,col2, op2, val2对应第二个列的匹配规则(多条件时用&逻辑合并)
  • 最后一列是匹配成功后的结果值
  • 标记default的行对应所有条件不匹配时的默认值

2. 读取配置并生成结果

import pandas as pd
import numpy as np

# 读取配置文件
config_df = pd.read_csv('config.csv', header=None, names=['col1', 'op1', 'val1', 'col2', 'op2', 'val2', 'result'])

conditions = []
choices = []
default_val = None

for _, row in config_df.iterrows():
    # 初始化条件
    current_cond = None
    
    # 处理第一个匹配规则
    if pd.notna(row['col1']) and pd.notna(row['op1']) and pd.notna(row['val1']):
        if row['op1'] == '==':
            current_cond = df[row['col1']] == row['val1']
        # 扩展子串匹配支持
        elif row['op1'] == 'contains':
            current_cond = df[row['col1']].str.contains(row['val1'], na=False)
    
    # 处理第二个匹配规则(多条件组合)
    if pd.notna(row['col2']) and pd.notna(row['op2']) and pd.notna(row['val2']):
        if row['op2'] == '==':
            sub_cond = df[row['col2']] == row['val2']
        elif row['op2'] == 'contains':
            sub_cond = df[row['col2']].str.contains(row['val2'], na=False)
        # 用&合并多条件
        current_cond = current_cond & sub_cond
    
    # 处理默认值
    if row['op1'] == 'default':
        default_val = row['result']
    else:
        conditions.append(current_cond)
        choices.append(row['result'])

# 用numpy.select批量填充目标列
df['target_col'] = np.select(conditions, choices, default=default_val)

方案2:使用YAML配置文件(适合复杂条件组合)

YAML格式更易读,支持更灵活的条件嵌套或多逻辑组合(如|或逻辑)。

1. 编写配置文件

创建config.yaml:

conditions:
  - cols:
      - [Set, ==, Z]
      - [Type, ==, A]
    result: yellow
  - cols:
      - [Set, ==, Z]
      - [Type, ==, B]
    result: blue
  - cols:
      - [Set, ==, X]
      - [Type, ==, B]
    result: purple
  - cols:
      - [Description, contains, "urgent"]  # 子串匹配示例
    result: red
default: black

2. 读取配置并生成结果

import pandas as pd
import numpy as np
import yaml

# 读取YAML配置
with open('config.yaml', 'r', encoding='utf-8') as f:
    config = yaml.safe_load(f)

conditions = []
choices = []

# 解析每个条件
for item in config['conditions']:
    current_cond = None
    for col_rule in item['cols']:
        col, op, val = col_rule
        # 生成单条件表达式
        if op == '==':
            sub_cond = df[col] == val
        elif op == 'contains':
            sub_cond = df[col].str.contains(val, na=False)
        # 合并多条件(默认用&,需或逻辑可替换为|)
        current_cond = sub_cond if current_cond is None else current_cond & sub_cond
    conditions.append(current_cond)
    choices.append(item['result'])

# 填充目标列
df['target_col'] = np.select(conditions, choices, default=config['default'])

方案优势

  • 维护便捷:新增/修改条件只需编辑配置文件,无需改动代码逻辑
  • 扩展性强:可轻松扩展支持!=、>等其他操作符,或|(或)逻辑组合
  • 可读性高:结构化配置比硬编码的条件列表更清晰,降低出错概率

内容的提问来源于stack exchange,提问作者cyphex

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.24 23:11:14