You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用通配符简化映射子串到单一值的Python字典?

解决方案:用正则匹配实现通配符字典映射

步骤1:定义简化的通配符字典

把重复的具体键替换为带通配符*的规则,*代表任意长度的任意字符:

# 简化后的通配符字典
col_dict = {
    'pur_act': 'ACQ1',
    '_depr': 'DEPR',
    'future_dated*': 'FDIO',
    'cap cost*': 'MANU',
    '*reclass_*': 'MANU',
    'funded*': 'MANU',
    'manual*': 'MANU',
    'non-tax*': 'MANU',
    '*other_*': 'MANU',
    'precomm*': 'MANU',
    '_int': 'OINT',
    '_additions': 'OPAA',
    '_pmt': 'OPMT',
    '_fx': 'OTFX',
    '*sub*': 'SUBS',
    'tia_*': 'TIOP',
    'unpaid*': 'UNIO'
}

通配符位置说明:

  • prefix*:匹配以prefix开头的字符串
  • *suffix:匹配以suffix结尾的字符串
  • *substr*:匹配包含substr的字符串
  • 无*:精确匹配字符串

步骤2:转换通配符为正则表达式

将通配符规则转成正则表达式(转义特殊字符,替换*为.*),并编译提升匹配效率:

import re

# 生成正则规则列表,按匹配精度从高到低排序(避免宽泛规则优先匹配)
regex_rules = []
for pattern, code in col_dict.items():
    # 转义正则特殊字符,再替换通配符*为正则的任意匹配
    regex_pattern = re.escape(pattern).replace(r'\*', '.*')
    regex_rules.append((re.compile(regex_pattern), code))

步骤3:自定义映射函数

写一个函数遍历正则规则,返回第一个匹配到的编码:

def map_column_name(col_name):
    for regex, code in regex_rules:
        # 用match做前缀匹配,改成search可实现包含匹配
        if regex.search(col_name):
            return code
    # 无匹配时返回默认值
    return 'UNKNOWN'

步骤4:应用到DataFrame

用apply方法将函数作用于column_name列,生成code列:

import pandas as pd

# 示例DataFrame
data = {
    'column_name': [
        'forecasted_rou_asset_additions',
        'forecasted_rou_liability_additions',
        'commenced_leases_fcst_depr',
        'forecasted_additions_fcst_depr',
        'yoy_fx',
        'commenced_leases_fcst_int',
        'forecasted_leases_fcst_int',
        'commenced_leases_fcst_pmt',
        'forecasted_leases_fcst_pmt',
        'tax amort cap cost_abc385',
        'tax amort cap cost_abc385',
        'funded const commit_abc385',
        'funded const commit_abc385',
        'future_dated_invoices',
        'manual_adjustment_fcst_abc385',
        'manual_adjustment_fcst_abc385',
        'non-tax amort cap cost_abc385'
    ]
}
df = pd.DataFrame(data)

# 生成code列
df['code'] = df['column_name'].apply(map_column_name)

验证结果

生成的code列与预期完全匹配:

column_namecode
forecasted_rou_asset_additionsOPAA
forecasted_rou_liability_additionsOPAA
commenced_leases_fcst_deprDEPR
forecasted_additions_fcst_deprDEPR
yoy_fxOTFX
commenced_leases_fcst_intOINT
forecasted_leases_fcst_intOINT
commenced_leases_fcst_pmtOPMT
forecasted_leases_fcst_pmtOPMT
tax amort cap cost_abc385MANU
tax amort cap cost_abc385MANU
funded const commit_abc385MANU
funded const commit_abc385MANU
future_dated_invoicesFDIO
manual_adjustment_fcst_abc385MANU
manual_adjustment_fcst_abc385MANU
non-tax amort cap cost_abc385MANU

简洁写法优化

如果规则顺序已按优先级排列,可直接用replace的正则模式实现:

# 转换为正则键的字典(Python3.7+保留插入顺序)
regex_dict = {re.escape(k).replace(r'\*', '.*'): v for k, v in col_dict.items()}
df['code'] = df['column_name'].replace(regex_dict, regex=True)

内容的提问来源于stack exchange,提问作者Jim

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.13 02:07:05