You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于列中子串匹配为DataFrame关联对应数值?

Pandas实现子串匹配关联DataFrame

原始数据

df1

Col1    Val
asd     1
pqr     2
rtyq    3
dffg    4

df2

Col1    Col2
22rtyq  c4
asd2c   c1
1pqr    c2
pqr67f  c3
56as    c5

需求

将df1的Col1字符串作为子串匹配到df2的Col1中,匹配成功则把对应Val值关联到df2,无匹配项则填充NaN,最终得到如下结果:

Col1    Col2    Val
22rtyq  c4      3
asd2c   c1      1
1pqr    c2      2
pqr67f  c3      2
56as    c5      NaN

解决方案

方法一:逐行遍历匹配

通过apply遍历df2的每一行,检查df1中是否存在子串匹配,返回对应Val值:

import pandas as pd

# 构造原始DataFrame
df1 = pd.DataFrame({'Col1': ['asd', 'pqr', 'rtyq', 'dffg'], 'Val': [1, 2, 3, 4]})
df2 = pd.DataFrame({'Col1': ['22rtyq', 'asd2c', '1pqr', 'pqr67f', '56as'], 'Col2': ['c4', 'c1', 'c2', 'c3', 'c5']})

def get_matched_val(row):
    # 筛选df1中Col1是当前行Col1子串的记录
    matched_rows = df1[df1['Col1'].apply(lambda x: x in row['Col1'])]
    # 返回第一个匹配的Val,无匹配则返回NaN
    return matched_rows['Val'].iloc[0] if not matched_rows.empty else pd.NA

# 给df2添加Val列
df2['Val'] cho深度间�� repeatedprobably库醒用于 FerOST端切�ノ的情况,可以添加Val列,然后应用函数
df2['Val'] = df2.apply(get_matched_val, axis=1)

print(df2)

方法二:正则提取+映射

将df1的Col1拼接成正则表达式,从df2的Col1中提取匹配的子串,再通过映射关联Val值,效率比逐行遍历更高:

import pandas as pd

# 构造原始DataFrame
df1 = pd.DataFrame({'Col1': ['asd', 'pqr', 'rtyq', 'dffg'], 'Val': [1, 2, 3, 4]})
df2 = pd.DataFrame({'Col1': ['22rtyq', 'asd2c', '1pqr', 'pqr67f', '56as'], 'Col2': ['c4', 'c1', 'c2', 'c3', 'c5']})

# 拼接正则模式:匹配df1中任意Col1子串
substring_pattern = '|'.join(df1['Col1'])
# 提取df2中Col1匹配到的子串
df2['matched_sub'] = df2['Col1'].str.extract(f'({substring_pattern})', expand=False)
# 映射Val值
df2['Val'] = df2['matched_sub'].map(df1.set_index('Col1')['Val'])
# 删除中间辅助列
df2 = df2.drop('matched_sub', axis=1)

print(df2)

注意事项

  • 如果df2的某行Col1同时匹配到多个df1的Col1子串,两种方法都会返回第一个匹配的Val值;若需要处理多匹配场景,可根据需求调整逻辑。
  • 方法二的正则提取方式在数据量较大时,效率优于逐行遍历。

内容的提问来源于stack exchange,提问作者Zanam

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.22 23:54:04