You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

查找Pandas DataFrame索引与列值配对是否存在的最快方法

性能优化方案

你当前的查询慢核心原因是所有方案都在做全表级的条件扫描,时间复杂度为O(n),数据量越大耗时越高。换成基于哈希结构的O(1)查找方案,性能可以提升上百倍:

方案1:重构为复合索引(推荐)

Pandas的索引本身是哈希表结构,直接把Date索引和Type列合并为复合索引,查找效率最高,也不需要额外维护其他数据结构:

# 首次初始化时执行一次,将Type列加入索引形成[Date, Type]的复合索引
compiledData = compiledData.set_index('Type', append=True)

# 后续检查配对是否存在的代码,耗时普遍低于1ms
existing_pair = (newData['Date'], newData['Type']) in compiledData.index

# 新增数据逻辑
if not existing_pair:
    # 把新数据也转为相同格式的复合索引再合并
    new_row = pd.DataFrame([newData]).set_index(['Date', 'Type'])
    compiledData = pd.concat([compiledData, new_row])

方案2:预存配对集合(不改动原DataFrame结构)

如果需要保留原单日期索引的结构,可以单独维护一个存储所有已存在配对的set,set的in查询也是O(1)复杂度:

# 首次初始化时执行一次,生成所有已存在的配对集合
existing_pairs = set(zip(compiledData.index, compiledData['Type']))

# 检查配对是否存在
existing_pair = (newData['Date'], newData['Type']) in existing_pairs

# 新增数据逻辑
if not existing_pair:
    compiledData = pd.concat([compiledData, pd.DataFrame([newData]).set_index('Date')])
    # 同步更新配对集合,避免后续重复生成
    existing_pairs.add((newData['Date'], newData['Type']))

批量新增优化

如果是一次性导入多条新数据,不要逐条检查,直接批量过滤效率更高:

# 假设newData是包含多条待新增数据的DataFrame,已经设置Date索引、包含Type列
new_pairs = pd.MultiIndex.from_arrays([newData.index, newData['Type']])
# 直接过滤出所有不存在的行
to_append = newData[~new_pairs.isin(compiledData.index)]
# 批量合并
compiledData = pd.concat([compiledData, to_append])

内容的提问来源于stack exchange,提问作者scima96

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.06 03:36:03