如何在Pandas中提取两个DataFrame列间的最长公共子串?
为DataFrame添加对应行的最长公共子串列
解决步骤
- 导入
difflib.SequenceMatcher和pandas模块 - 封装单个字符串对的最长公共子串提取逻辑为函数
- 遍历两个DataFrame的对应行,批量生成目标列
完整代码示例
import pandas as pd from difflib import SequenceMatcher # 定义提取最长公共子串的函数 def get_longest_common_substring(s1, s2): match = SequenceMatcher(None, s1, s2).find_longest_match(0, len(s1), 0, len(s2)) # 处理无匹配的边界情况,返回空字符串 return s1[match.a: match.a + match.size] if match.size > 0 else "" # 创建示例DataFrame df_a = pd.DataFrame({"String": ["012IREze", "SecondString", "LastEntry"]}) df_b = pd.DataFrame({"String": ["IREPP", "StringNumber2", "LastEntry123"]}) # 新增目标列:逐对处理两个DataFrame的对应行字符串 df_a["Common String"] = [get_longest_common_substring(s1, s2) for s1, s2 in zip(df_a["String"], df_b["String"])] print(df_a)
输出结果
String Common String 0 012IREze IRE 1 SecondString String 2 LastEntry LastEntry
关键说明
- 函数
get_longest_common_substring复用了你提供的SequenceMatcher逻辑,新增了无匹配时的容错处理 - 使用
zip同步遍历两个DataFrame的String列,确保每行字符串对应匹配 - 最终直接在df_a中生成目标列,完全符合你给出的示例需求
内容的提问来源于stack exchange,提问作者PythonBeginner
相关产品推荐
相关产品推荐

