You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas:当DataFrame列匹配列表字符串时替换另一列值的方法

高效替换DataFrame中匹配条件的列值

嘿,嵌套循环确实又丑又慢,尤其是数据量上去之后完全顶不住!咱们用pandas的矢量化字符串操作来解决这个问题,既简洁又高效,完全不用写循环。

先把你的示例数据贴出来方便验证:

import pandas as pd

df = pd.DataFrame({
    'col1': ['str9', 'str8', 'str4', 'str2', 'unknown1', 'random1'],
    'col2': ['val1', 'val2', 'val3', 'val4', 'asdvstr1df', 'teststr2test']
})
strings = ['str1', 'str2', 'str3', 'str4']

核心思路

  1. 把你的匹配列表转换成正则表达式模式,用|分隔所有元素,这样可以一次性匹配col2中任意位置的目标字符串;
  2. 用pandas的str.contains快速标记所有符合条件的行(矢量化操作,比循环快N倍);
  3. 用str.extract提取col2中匹配到的具体字符串,替换col1对应的值。

完整实现代码

import numpy as np

# 构建正则模式:匹配任意位置的目标字符串
match_pattern = '|'.join(strings)

# 生成匹配掩码:col2中包含任意目标字符串的行标记为True
match_mask = df['col2'].str.contains(match_pattern, regex=True)

# 替换col1:匹配行用提取到的目标字符串替换,未匹配行保持原值
df['col1'] = np.where(
    match_mask,
    df['col2'].str.extract(f'({match_pattern})', expand=False),
    df['col1']
)

运行结果

执行后你的DataFrame会变成这样:

col1          col2
0  str9          val1
1  str8          val2
2  str4          val3
3  str2          val4
4  str1  asdvstr1df
5  str2  teststr2test

额外说明

如果col2中可能存在多个匹配字符串(比如同时包含str1和str2),可以用str.findall提取所有匹配项,再根据需求处理(比如取第一个、合并等):

# 提取所有匹配项并取第一个
df['col1'] = np.where(
    match_mask,
    df['col2'].str.findall(match_pattern).str[0],
    df['col1']
)

这种方法完全利用了pandas的内部优化,没有任何循环,数据量越大,效率提升越明显,代码也更易维护~

内容的提问来源于stack exchange,提问作者dward4

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 06:42:28