You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在DataFrame中匹配字符串子串并生成对应匹配列?

问题描述

我有如下DataFrame:

phraselen
i love2
he plays2
i love people3
love1

需求:为每个phrase单元格,查找其在其他phrase单元格中的出现情况,将结果展示在单独的match列中。

我尝试了以下代码:

for i in range(len(df['phrase'])):
     for j in range(len(df['phrase'])):
         if (df['phrase'].iloc[i] in df['phrase'].iloc[j]) and (df['phrase'].iloc[i] != df['phrase'].iloc[j]):
             df['match'].iloc[i]=df['phrase'].iloc[j]

预期输出如下:

phrasematch
i lovei love people
i love peoplenon
lovei love
lovei love people
he playsnon
解决方法

你的代码存在两个核心问题:一是双重循环遍历效率低,且直接用iloc赋值容易触发SettingWithCopyWarning;二是当一个短语匹配到多个结果时,只会保留最后一个匹配项,无法生成预期的多行结果。

下面是两种高效且符合预期的实现方式:

方法1:列表推导式(简洁直观)

import pandas as pd

# 构造原始DataFrame
df = pd.DataFrame({
    'phrase': ['i love', 'he plays', 'i love people', 'love'],
    'len': [2, 2, 3, 1]
})

phrases = df['phrase'].tolist()
matches_list = []

for phrase in phrases:
    # 收集所有包含当前短语的其他短语
    matched = [p for p in phrases if phrase in p and phrase != p]
    if matched:
        # 每个匹配项生成一行
        for p in matched:
            matches_list.append({'phrase': phrase, 'match': p})
    else:
        matches_list.append({'phrase': phrase, 'match': 'non'})

final_df = pd.DataFrame(matches_list)
print(final_df)

方法2:Pandas合并操作(适合大规模数据)

import pandas as pd

df = pd.DataFrame({
    'phrase': ['i love', 'he plays', 'i love people', 'love'],
    'len': [2, 2, 3, 1]
})

# 生成所有短语组合,排除自身匹配
cross_df = df.assign(key=1).merge(df.assign(key=1), on='key', suffixes=('', '_other'))
cross_df = cross_df[cross_df['phrase'] != cross_df['phrase_other']]

# 筛选出短语被包含的行
matched_df = cross_df[cross_df.apply(lambda x: x['phrase'] in x['phrase_other'], axis=1)]
matched_df = matched_df[['phrase', 'phrase_other']].rename(columns={'phrase_other': 'match'})

# 补充无匹配的短语
no_match_df = df[~df['phrase'].isin(matched_df['phrase'])]['phrase'].to_frame()
no_match_df['match'] = 'non'

# 合并结果
final_df = pd.concat([matched_df, no_match_df], ignore_index=True)
print(final_df)

两种方法都能输出符合预期的结果,可根据数据规模选择使用。


内容的提问来源于stack exchange,提问作者Михаил Беляков

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.03 11:20:37