You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Python中为基因DataFrame添加关联的GO ID注释列?

实现方法

核心思路

先拆分Table2中每个GO对应的多geneID,建立gene与GO信息的一对一映射,再聚合每个gene的所有GO注释内容,最后合并到Table1中生成目标结果。

步骤代码(基于Pandas)

1. 准备示例数据(替换为你的真实数据即可)

import pandas as pd

# 示例Table1:含唯一geneID及其他业务列
Table1 = pd.DataFrame({
    'geneID': ['geneA', 'geneB', 'geneC'],
    'expression': [2.5, 1.8, 3.2],
    'sample': ['Sample1', 'Sample2', 'Sample3']
})

# 示例Table2:含GO_ID、描述及斜杠分隔的关联geneID集合
Table2 = pd.DataFrame({
    'GO_ID': ['GO:0007049', 'GO:0007165', 'GO:0006810'],
    'description': ['Cell cycle', 'Signal transduction', 'Transport'],
    'geneIDs': ['geneA/geneB', 'geneB/geneC', 'geneA']
})

2. 拆分并构建gene-GO映射关系

# 拆分Table2的geneIDs列,将每个GO对应的多个gene拆分为单独行
table2_expanded = Table2.assign(geneID=Table2['geneIDs'].str.split('/')).explode('geneID')

# 组合GO_ID与描述为目标格式:"GO_ID 描述"
table2_expanded['GO_annotation'] = table2_expanded['GO_ID'] + ' ' + table2_expanded['description']

3. 聚合每个gene的所有GO注释

# 按geneID分组,将同一gene的所有GO注释用分号分隔拼接
gene_go_annotations = table2_expanded.groupby('geneID')['GO_annotation'].agg('; '.join).reset_index()

4. 合并到Table1生成最终结果

# 与原Table1合并,保留所有geneID(无GO注释的会显示NaN)
Table3 = Table1.merge(gene_go_annotations, on='geneID', how='left')

# 可选:将无注释的gene对应的GO列设为空字符串
Table3['GO_annotation'] = Table3['GO_annotation'].fillna('')

最终示例效果(Table3)

geneIDexpressionsampleGO_annotation
geneA2.5Sample1GO:0007049 Cell cycle; GO:0006810 Transport
geneB1.8Sample2GO:0007049 Cell cycle; GO:0007165 Signal transduction
geneC3.2Sample3GO:0007165 Signal transduction

内容的提问来源于stack exchange,提问作者Antuan Rage

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.07 22:00:23