You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Pandas中提取两个DataFrame列间的最长公共子串?

为DataFrame添加对应行的最长公共子串列

解决步骤

  • 导入difflib.SequenceMatcher和pandas模块
  • 封装单个字符串对的最长公共子串提取逻辑为函数
  • 遍历两个DataFrame的对应行,批量生成目标列

完整代码示例

import pandas as pd
from difflib import SequenceMatcher

# 定义提取最长公共子串的函数
def get_longest_common_substring(s1, s2):
    match = SequenceMatcher(None, s1, s2).find_longest_match(0, len(s1), 0, len(s2))
    # 处理无匹配的边界情况,返回空字符串
    return s1[match.a: match.a + match.size] if match.size > 0 else ""

# 创建示例DataFrame
df_a = pd.DataFrame({"String": ["012IREze", "SecondString", "LastEntry"]})
df_b = pd.DataFrame({"String": ["IREPP", "StringNumber2", "LastEntry123"]})

# 新增目标列:逐对处理两个DataFrame的对应行字符串
df_a["Common String"] = [get_longest_common_substring(s1, s2) for s1, s2 in zip(df_a["String"], df_b["String"])]

print(df_a)

输出结果

String Common String
0     012IREze           IRE
1  SecondString         String
2    LastEntry      LastEntry

关键说明

  • 函数get_longest_common_substring复用了你提供的SequenceMatcher逻辑,新增了无匹配时的容错处理
  • 使用zip同步遍历两个DataFrame的String列,确保每行字符串对应匹配
  • 最终直接在df_a中生成目标列,完全符合你给出的示例需求

内容的提问来源于stack exchange,提问作者PythonBeginner

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.22 14:15:35