查找Pandas DataFrame索引与列值配对是否存在的最快方法
性能优化方案
你当前的查询慢核心原因是所有方案都在做全表级的条件扫描,时间复杂度为O(n),数据量越大耗时越高。换成基于哈希结构的O(1)查找方案,性能可以提升上百倍:
方案1:重构为复合索引(推荐)
Pandas的索引本身是哈希表结构,直接把Date索引和Type列合并为复合索引,查找效率最高,也不需要额外维护其他数据结构:
# 首次初始化时执行一次,将Type列加入索引形成[Date, Type]的复合索引 compiledData = compiledData.set_index('Type', append=True) # 后续检查配对是否存在的代码,耗时普遍低于1ms existing_pair = (newData['Date'], newData['Type']) in compiledData.index # 新增数据逻辑 if not existing_pair: # 把新数据也转为相同格式的复合索引再合并 new_row = pd.DataFrame([newData]).set_index(['Date', 'Type']) compiledData = pd.concat([compiledData, new_row])
方案2:预存配对集合(不改动原DataFrame结构)
如果需要保留原单日期索引的结构,可以单独维护一个存储所有已存在配对的set,set的in查询也是O(1)复杂度:
# 首次初始化时执行一次,生成所有已存在的配对集合 existing_pairs = set(zip(compiledData.index, compiledData['Type'])) # 检查配对是否存在 existing_pair = (newData['Date'], newData['Type']) in existing_pairs # 新增数据逻辑 if not existing_pair: compiledData = pd.concat([compiledData, pd.DataFrame([newData]).set_index('Date')]) # 同步更新配对集合,避免后续重复生成 existing_pairs.add((newData['Date'], newData['Type']))
批量新增优化
如果是一次性导入多条新数据,不要逐条检查,直接批量过滤效率更高:
# 假设newData是包含多条待新增数据的DataFrame,已经设置Date索引、包含Type列 new_pairs = pd.MultiIndex.from_arrays([newData.index, newData['Type']]) # 直接过滤出所有不存在的行 to_append = newData[~new_pairs.isin(compiledData.index)] # 批量合并 compiledData = pd.concat([compiledData, to_append])
内容的提问来源于stack exchange,提问作者scima96
相关产品推荐
相关产品推荐

