You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用字符串正则模式映射pandas Series时返回NaN如何解决

Pandas使用正则规则映射Series列的实现方案

原有map方法失效的原因是:Series.map()传入字典时执行精确字符串匹配,不会自动将字典键识别为正则表达式做模糊匹配,因此全部匹配失败返回NaN。

以下是可直接运行的实现方案:

方案1:使用numpy.select实现(多规则场景优先推荐)

该方法扩展性强,新增正则规则只需对应添加条件和返回值即可:

import pandas as pd
import numpy as np

# 示例数据
s = pd.DataFrame([['AMcU8', 10], ['AM8v', 15], ['ASw9', 14],['ASw7', 14]], columns = ['Code', 'Quantity'])

# 定义正则匹配条件列表
conditions = [
    s['Code'].str.match(r'AM.*8.*'),
    s['Code'].str.match(r'AS.*9.*')
]
# 定义匹配成功后对应的返回值列表,顺序和conditions一一对应
values = ['AM8', 'AS9']

# 匹配不到的项默认返回NaN
s['newcode'] = np.select(conditions, values, default=np.nan)

运行后newcode列的输出结果为:['AM8', 'AM8', 'AS9', NaN],完全符合需求。

方案2:使用str.replace实现(规则简单场景适用)

如果正则规则较少,也可以通过链式调用str.replace实现:

import pandas as pd
import numpy as np

s = pd.DataFrame([['AMcU8', 10], ['AM8v', 15], ['ASw9', 14],['ASw7', 14]], columns = ['Code', 'Quantity'])

# 依次按正则规则替换匹配到的值
s['newcode'] = s['Code'].str.replace(r'^AM.*8.*$', 'AM8', regex=True)\
                        .str.replace(r'^AS.*9.*$', 'AS9', regex=True)
# 未匹配到的原值转为NaN
s['newcode'] = s['newcode'].where(s['newcode'].isin(['AM8', 'AS9']), np.nan)

内容的提问来源于stack exchange,提问作者Alberto B

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.29 08:54:06