You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何高效匹配两个DataFrame中的电话号码并生成匹配列?

高效实现电话号码匹配方案

原代码采用双重循环遍历两个DataFrame,时间复杂度为O(n*m),数据量较大时执行效率极低。以下是基于集合查询的高效优化方案:

优化思路

  • 将df2的电话号码存入集合:集合的成员查询操作时间复杂度为O(1),远快于遍历DataFrame的O(m)
  • 对df1的每一行号码进行拆分,直接和集合求交集,彻底避免双重循环

具体实现代码

import pandas as pd

# 预处理:将df2的电话号码转为字符串并存入集合(去重进一步提升查询效率)
df2['telephone'] = df2['telephone'].astype(str)
df2_tel_set = set(df2['telephone'].unique())

# 定义匹配函数:拆分号码并查找匹配项
def get_matching_tel(tel_str):
    # 拆分逗号分隔的号码,同时去除可能存在的空格
    split_tels = [tel.strip() for tel in tel_str.split(',')]
    # 筛选出在df2集合中的号码
    matched = [tel for tel in split_tels if tel in df2_tel_set]
    # 与原逻辑一致:返回第一个匹配项,无匹配则返回False
    return matched[0] if matched else False

# 应用到df1,生成matches列
df1['telephone'] = df1['telephone'].astype(str)
df1['matches'] = df1['telephone'].apply(get_matching_tel)

扩展说明

如果需要返回所有匹配的电话号码(而非仅第一个),可以修改函数的返回逻辑:

def get_all_matching_tels(tel_str):
    split_tels = [tel.strip() for tel in tel_str.split(',')]
    matched = [tel for tel in split_tels if tel in df2_tel_set]
    return ','.join(matched) if matched else False

效率对比

原方案的时间复杂度为O(n*m)(n为df1行数,m为df2行数),优化后的方案时间复杂度为O(n + m),当df2数据量达到万级以上时,执行速度会有数量级的提升。

内容的提问来源于stack exchange,提问作者user3347814

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.22 06:31:10