You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何更快在terms数据框中检索df的gene1-gene2基因对?

高效检测基因对是否为基因集合子集的优化方案

数据背景

  • 大型DataFrame df:约4500万行,结构如下:
gene1   gene2     score
0       PIGA  ATF7IP1 -0.047236
1       PIGB  ATF7IP2 -0.047236
2       PIGC  ATF7IP3 -0.047236
3       PIGD  ATF7IP4 -0.047236
4       PIGE  ATF7IP5 -0.047236
  • 小型DataFrame terms:约3000行,结构如下:
id                                gene_set
1                                 {HDAC4, BCL6}
2                                 {HDAC5, BCL6}
3                                 {HDAC7, BCL6}
4                 {NCOA3, KAT2B, EP300, CREBBP}
5            {NCAPD2, NCAPH, NCAPG, SMC4, SMC2}
                         ...                   
2912                              {FOXO1, ESR1}
2913                               {APP, FOXO3}
2914                               {APP, FOXO1}
2915                               {APP, FOXO4}
2916    {MAP3K20, MAPK14, AKAP13, MAP2K3, PKN1}

需求

对df的每一行,检查gene1与gene2组成的基因对是否是terms中某个gene_set的子集,统计每个基因对匹配到的gene_set数量。

现有低效实现

目前的几种实现效率相近,但面对4500万行数据速度极慢:

搜索函数定义

def search(g1,g2):
    # 搜索基因对是否在基因集合中
    return sum(terms.gene_set.map(set([g1,g2]).issubset))

调用方式

  1. 向量式调用:
np.sum(np.vectorize(search)(df.gene1,df.gene2))
  1. 列表推导:
[search(g1, g2) for g1, g2 in zip(df.gene1,df.gene2)] 
  1. apply方法:
df[['gene1','gene2']].apply(lambda x: search(x.gene1,x.gene2), axis=1 )

高效优化思路

核心逻辑是反转匹配方向:不遍历4500万行基因对去匹配3000个集合,而是先遍历3000个集合生成所有有效二元基因对,用哈希表存储匹配计数,再批量查询df中的基因对。

步骤1:预处理terms生成基因对-计数映射

from collections import defaultdict
import itertools

# 构建字典:键为排序后的基因对元组,值为匹配的gene_set数量
pair_to_count = defaultdict(int)

for gene_set in terms['gene_set']:
    # 仅处理包含至少2个基因的集合
    if len(gene_set) >= 2:
        # 生成集合内所有无序二元对,排序后避免重复(如(BCL6, HDAC4)与(HDAC4, BCL6)视为同一对)
        for pair in itertools.combinations(sorted(gene_set), 2):
            pair_to_count[pair] += 1

步骤2:批量处理df中的基因对

# 快速查询函数
def get_count(g1, g2):
    sorted_pair = tuple(sorted((g1, g2)))
    return pair_to_count.get(sorted_pair, 0)

# 批量计算每个基因对的匹配数
df['match_count'] = df.apply(lambda row: get_count(row['gene1'], row['gene2']), axis=1)

# 若需要总匹配数,直接求和
total_matches = df['match_count'].sum()

优化原理

  • 时间复杂度从O(4500万 × 3000) 降至O(3000 × K + 4500万),其中K为每个gene_set的平均基因数(远小于3000)。
  • 哈希表查询为O(1)操作,彻底避免了每次遍历整个terms表的重复计算。
  • 预处理仅需执行一次,后续可重复使用映射结果。

额外优化建议

  • 统一基因名称格式(如大小写、去除特殊字符),避免因格式不一致导致匹配失败。
  • 用矢量化操作替代apply进一步提速:
# 生成排序后的基因对列
df['sorted_pair'] = df.apply(lambda row: tuple(sorted((row['gene1'], row['gene2']))), axis=1)
# 快速映射计数
df['match_count'] = df['sorted_pair'].map(pair_to_count).fillna(0).astype(int)
  • 若内存充足,可将pair_to_count转为pandas.Series,提升映射效率。

内容的提问来源于stack exchange,提问作者Yasir

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.25 13:48:20