You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Python中根据列值使用不同正则规则条件提取字符串

在Python中根据列值匹配不同正则提取字符串

需求说明

给定如下DataFrame:

col1   col2
1          John Smith  First
2          Jane Smith  First
3 Pritchard James Doe Second
4    Helen Joanne Doe Second
5         Walker Jean   Last
6         Hall Jensen   Last

期望提取后得到:

col1   col2   col3
1          John Smith  First   John
2          Jane Smith  First   Jane
3 Pritchard James Doe Second  James
4    Helen Joanne Doe Second Joanne
5         Walker Jean   Last   Jean
6         Hall Jensen   Last Jensen

在R中可以用case_when结合str_extract实现:

library(tidyverse)
df <- data.frame(col1 = c("John Smith", "Jane Smith", "Pritchard James Doe", "Helen Joanne Doe", "Walker Jean", "Hall Jensen"),
                 col2 = c("First", "First", "Second", "Second", "Last", "Last"))
df %>%
    mutate(col3 = case_when(col2 == "First" ~ str_extract(col1, "^[A-Za-z]+"),
                            col2 == "Second" ~ str_extract(col1, "(?<=\\s+)[A-Za-z]+"),
                            col2 == "Last" ~ str_extract(col1, "[A-Za-z]+$")))

但在Python的pandas中尝试用case_when结合lambda和正则未成功,尝试代码如下:

import pandas as pd
import re
d = {'col1': ['John Smith', 'Jane Smith', 'Pritchard James Doe', 'Helen Joanne Doe', 'Walker Jean', 'Hall Jensen'],
     'col2': ['First', 'First', 'Second', 'Second', 'Last', 'Last']}
df = pd.DataFrame(d)
cl = [(df['col2'] == 'First', lambda x: re.search('(^[A-Za-z]+)', x).group()),
      (df['col2'] == 'Second', lambda x: re.search(r'(?<=\\s)([A-Za-z]+)', x).group()),
      (df['col2'] == 'Last', lambda x: re.search('([A-Za-z]+$)', x).group())]
df.assign(col3 = df['col1'].case_when(cl))

解决方案

方法一:使用numpy.select(最贴近R的case_when逻辑)

通过定义条件列表和对应的提取规则,批量完成匹配:

import pandas as pd
import numpy as np

d = {'col1': ['John Smith', 'Jane Smith', 'Pritchard James Doe', 'Helen Joanne Doe', 'Walker Jean', 'Hall Jensen'],
     'col2': ['First', 'First', 'Second', 'Second', 'Last', 'Last']}
df = pd.DataFrame(d)

# 定义匹配条件
conditions = [
    df['col2'] == 'First',
    df['col2'] == 'Second',
    df['col2'] == 'Last'
]

# 对应条件的正则提取规则
extractors = [
    df['col1'].str.extract(r'^([A-Za-z]+)', expand=False),
    df['col1'].str.extract(r'\s+([A-Za-z]+)', expand=False),
    df['col1'].str.extract(r'([A-Za-z]+)$', expand=False)
]

# 应用条件选择生成col3
df['col3'] = np.select(conditions, extractors)
print(df)

方法二:使用apply逐行处理

适合逻辑更复杂的场景,逐行判断并提取:

import pandas as pd
import re

d = {'col1': ['John Smith', 'Jane Smith', 'Pritchard James Doe', 'Helen Joanne Doe', 'Walker Jean', 'Hall Jensen'],
     'col2': ['First', 'First', 'Second', 'Second', 'Last', 'Last']}
df = pd.DataFrame(d)

def extract_target(row):
    if row['col2'] == 'First':
        return re.search(r'^[A-Za-z]+', row['col1']).group()
    elif row['col2'] == 'Second':
        return re.search(r'(?<=\s)[A-Za-z]+', row['col1']).group()
    elif row['col2'] == 'Last':
        return re.search(r'[A-Za-z]+$', row['col1']).group()

df['col3'] = df.apply(extract_target, axis=1)
print(df)

方法三:分批次用loc匹配处理

针对每个条件单独提取,再合并结果:

import pandas as pd

d = {'col1': ['John Smith', 'Jane Smith', 'Pritchard James Doe', 'Helen Joanne Doe', 'Walker Jean', 'Hall Jensen'],
     'col2': ['First', 'First', 'Second', 'Second', 'Last', 'Last']}
df = pd.DataFrame(d)

# 初始化col3
df['col3'] = ''

# 按条件分别提取
df.loc[df['col2'] == 'First', 'col3'] = df.loc[df['col2'] == 'First', 'col1'].str.extract(r'^([A-Za-z]+)', expand=False)
df.loc[df['col2'] == 'Second', 'col3'] = df.loc[df['col2'] == 'Second', 'col1'].str.extract(r'\s+([A-Za-z]+)', expand=False)
df.loc[df['col2'] == 'Last', 'col3'] = df.loc[df['col2'] == 'Last', 'col1'].str.extract(r'([A-Za-z]+)$', expand=False)

print(df)

原尝试失败原因

原生pandas的Series并没有case_when方法(该方法属于第三方库pyjanitor),直接调用会导致报错,以上几种方式都是原生pandas的可行替代方案。


内容的提问来源于stack exchange,提问作者combperm

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.18 11:33:19