如何在DataFrame中匹配字符串子串并生成对应匹配列?
问题描述
我有如下DataFrame:
| phrase | len |
|---|---|
| i love | 2 |
| he plays | 2 |
| i love people | 3 |
| love | 1 |
需求:为每个phrase单元格,查找其在其他phrase单元格中的出现情况,将结果展示在单独的match列中。
我尝试了以下代码:
for i in range(len(df['phrase'])): for j in range(len(df['phrase'])): if (df['phrase'].iloc[i] in df['phrase'].iloc[j]) and (df['phrase'].iloc[i] != df['phrase'].iloc[j]): df['match'].iloc[i]=df['phrase'].iloc[j]
预期输出如下:
| phrase | match |
|---|---|
| i love | i love people |
| i love people | non |
| love | i love |
| love | i love people |
| he plays | non |
解决方法
你的代码存在两个核心问题:一是双重循环遍历效率低,且直接用iloc赋值容易触发SettingWithCopyWarning;二是当一个短语匹配到多个结果时,只会保留最后一个匹配项,无法生成预期的多行结果。
下面是两种高效且符合预期的实现方式:
方法1:列表推导式(简洁直观)
import pandas as pd # 构造原始DataFrame df = pd.DataFrame({ 'phrase': ['i love', 'he plays', 'i love people', 'love'], 'len': [2, 2, 3, 1] }) phrases = df['phrase'].tolist() matches_list = [] for phrase in phrases: # 收集所有包含当前短语的其他短语 matched = [p for p in phrases if phrase in p and phrase != p] if matched: # 每个匹配项生成一行 for p in matched: matches_list.append({'phrase': phrase, 'match': p}) else: matches_list.append({'phrase': phrase, 'match': 'non'}) final_df = pd.DataFrame(matches_list) print(final_df)
方法2:Pandas合并操作(适合大规模数据)
import pandas as pd df = pd.DataFrame({ 'phrase': ['i love', 'he plays', 'i love people', 'love'], 'len': [2, 2, 3, 1] }) # 生成所有短语组合,排除自身匹配 cross_df = df.assign(key=1).merge(df.assign(key=1), on='key', suffixes=('', '_other')) cross_df = cross_df[cross_df['phrase'] != cross_df['phrase_other']] # 筛选出短语被包含的行 matched_df = cross_df[cross_df.apply(lambda x: x['phrase'] in x['phrase_other'], axis=1)] matched_df = matched_df[['phrase', 'phrase_other']].rename(columns={'phrase_other': 'match'}) # 补充无匹配的短语 no_match_df = df[~df['phrase'].isin(matched_df['phrase'])]['phrase'].to_frame() no_match_df['match'] = 'non' # 合并结果 final_df = pd.concat([matched_df, no_match_df], ignore_index=True) print(final_df)
两种方法都能输出符合预期的结果,可根据数据规模选择使用。
内容的提问来源于stack exchange,提问作者Михаил Беляков
相关产品推荐
相关产品推荐

