You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python提取GO层级Level2 ID匹配注释GO ID问题求助

修复GO术语Level 2 ID提取错误的Python方案

问题描述

我编写Python脚本时遇到以下问题:

  • 需匹配file.txt中注释的GO ID
  • 从go-basic.obo的层级结构提取Level 2 GO ID
  • 最终生成包含GO ID、术语名、Level 2 GO ID、频率的Excel文件

但现有代码无法正确提取目标Level 2 ID,例如GO:0006996的Level 2 ID被错误提取为Level 1的GO:0016043。

问题根源

错误原因是代码仅获取了GO term的直接父节点(即Level 1),未按照GO层级的定义(根节点为Level 1,其直接子节点为Level 2,下属节点需追溯到对应Level 2祖先)进行递归遍历。GO本体是有向无环图,每个term可能有多个父节点,需明确沿is_a关系追溯到三大根节点(生物过程GO:0008150、细胞组分GO:0005575、分子功能GO:0003674),再定位到根节点的直接子节点(即Level 2)。

修复后的代码

import pandas as pd
from collections import defaultdict

# 解析go-basic.obo文件,构建GO term的属性和父节点链
def parse_obo(obo_path):
    go_data = {}
    current_term = None
    with open(obo_path, 'r', encoding='utf-8') as f:
        for line in f:
            line = line.strip()
            if line == '[Term]':
                if current_term:
                    go_data[current_term['id']] = current_term
                current_term = {'id': '', 'name': '', 'is_a': []}
            elif line.startswith('id:'):
                current_term['id'] = line.split(': ')[1]
            elif line.startswith('name:'):
                current_term['name'] = line.split(': ')[1]
            elif line.startswith('is_a:'):
                # 提取父节点ID,忽略注释部分
                parent_id = line.split(' ! ')[0].split(': ')[1]
                current_term['is_a'].append(parent_id)
        # 处理最后一个term
        if current_term:
            go_data[current_term['id']] = current_term
    return go_data

# 定义GO三大根节点
ROOT_GOS = {'GO:0008150', 'GO:0005575', 'GO:0003674'}

# 获取目标GO对应的Level 2祖先节点
def get_level2_go(go_id, go_data):
    # 追溯到根节点,记录路径
    path = []
    current = go_id
    while current not in ROOT_GOS:
        if current not in go_data or not go_data[current]['is_a']:
            # 无父节点或不存在的GO ID,返回空
            return None
        path.append(current)
        # 取第一个is_a父节点(多父节点场景可根据需求调整逻辑)
        current = go_data[current]['is_a'][0]
    path.append(current)
    
    # 反转路径,得到从根到目标GO的顺序
    path.reverse()
    # 根为Level 1,Level 2是根的直接子节点(路径索引1)
    if len(path) >= 2:
        return path[1]
    # 若GO本身是根节点,无Level 2
    return None

# 统计file.txt中GO ID的出现频率
def count_go_occurrences(txt_path):
    go_counts = defaultdict(int)
    with open(txt_path, 'r', encoding='utf-8') as f:
        for line in f:
            go_id = line.strip()
            if go_id.startswith('GO:'):
                go_counts[go_id] += 1
    return go_counts

# 生成结果Excel
def generate_excel(go_data, go_counts, output_path):
    rows = []
    for go_id, freq in go_counts.items():
        if go_id not in go_data:
            continue
        term_name = go_data[go_id]['name']
        level2_id = get_level2_go(go_id, go_data)
        level2_name = go_data[level2_id]['name'] if level2_id else '无匹配Level 2节点'
        rows.append({
            'GO ID': go_id,
            '术语名': term_name,
            'Level 2 GO ID': level2_id,
            'Level 2 术语名': level2_name,
            '频率': freq
        })
    
    df = pd.DataFrame(rows)
    df.to_excel(output_path, index=False)
    print(f"结果已导出至 {output_path}")

if __name__ == '__main__':
    # 替换为你的文件路径
    go_obo_path = 'go-basic.obo'
    go_txt_path = 'file.txt'
    output_excel = 'go_level2_statistics.xlsx'
    
    go_terms = parse_obo(go_obo_path)
    go_freq = count_go_occurrences(go_txt_path)
    generate_excel(go_terms, go_freq, output_excel)

关键修复说明

  • 明确层级定义:以GO三大根节点为Level 1,其直接子节点为Level 2,确保所有下属GO term都映射到对应的Level 2分类
  • 完整路径追溯:遍历is_a关系追溯到根节点,通过路径反转准确定位Level 2节点,而非仅取直接父节点
  • 异常处理:对不存在的GO ID或无父节点的情况返回明确标识,避免程序崩溃
  • 灵活适配:多父节点场景可调整父节点选择逻辑(当前取第一个is_a父节点)

内容的提问来源于stack exchange,提问作者Umar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.13 10:52:32