Python提取GO层级Level2 ID匹配注释GO ID问题求助
修复GO术语Level 2 ID提取错误的Python方案
问题描述
我编写Python脚本时遇到以下问题:
- 需匹配
file.txt中注释的GO ID - 从
go-basic.obo的层级结构提取Level 2 GO ID - 最终生成包含GO ID、术语名、Level 2 GO ID、频率的Excel文件
但现有代码无法正确提取目标Level 2 ID,例如GO:0006996的Level 2 ID被错误提取为Level 1的GO:0016043。
问题根源
错误原因是代码仅获取了GO term的直接父节点(即Level 1),未按照GO层级的定义(根节点为Level 1,其直接子节点为Level 2,下属节点需追溯到对应Level 2祖先)进行递归遍历。GO本体是有向无环图,每个term可能有多个父节点,需明确沿is_a关系追溯到三大根节点(生物过程GO:0008150、细胞组分GO:0005575、分子功能GO:0003674),再定位到根节点的直接子节点(即Level 2)。
修复后的代码
import pandas as pd from collections import defaultdict # 解析go-basic.obo文件,构建GO term的属性和父节点链 def parse_obo(obo_path): go_data = {} current_term = None with open(obo_path, 'r', encoding='utf-8') as f: for line in f: line = line.strip() if line == '[Term]': if current_term: go_data[current_term['id']] = current_term current_term = {'id': '', 'name': '', 'is_a': []} elif line.startswith('id:'): current_term['id'] = line.split(': ')[1] elif line.startswith('name:'): current_term['name'] = line.split(': ')[1] elif line.startswith('is_a:'): # 提取父节点ID,忽略注释部分 parent_id = line.split(' ! ')[0].split(': ')[1] current_term['is_a'].append(parent_id) # 处理最后一个term if current_term: go_data[current_term['id']] = current_term return go_data # 定义GO三大根节点 ROOT_GOS = {'GO:0008150', 'GO:0005575', 'GO:0003674'} # 获取目标GO对应的Level 2祖先节点 def get_level2_go(go_id, go_data): # 追溯到根节点,记录路径 path = [] current = go_id while current not in ROOT_GOS: if current not in go_data or not go_data[current]['is_a']: # 无父节点或不存在的GO ID,返回空 return None path.append(current) # 取第一个is_a父节点(多父节点场景可根据需求调整逻辑) current = go_data[current]['is_a'][0] path.append(current) # 反转路径,得到从根到目标GO的顺序 path.reverse() # 根为Level 1,Level 2是根的直接子节点(路径索引1) if len(path) >= 2: return path[1] # 若GO本身是根节点,无Level 2 return None # 统计file.txt中GO ID的出现频率 def count_go_occurrences(txt_path): go_counts = defaultdict(int) with open(txt_path, 'r', encoding='utf-8') as f: for line in f: go_id = line.strip() if go_id.startswith('GO:'): go_counts[go_id] += 1 return go_counts # 生成结果Excel def generate_excel(go_data, go_counts, output_path): rows = [] for go_id, freq in go_counts.items(): if go_id not in go_data: continue term_name = go_data[go_id]['name'] level2_id = get_level2_go(go_id, go_data) level2_name = go_data[level2_id]['name'] if level2_id else '无匹配Level 2节点' rows.append({ 'GO ID': go_id, '术语名': term_name, 'Level 2 GO ID': level2_id, 'Level 2 术语名': level2_name, '频率': freq }) df = pd.DataFrame(rows) df.to_excel(output_path, index=False) print(f"结果已导出至 {output_path}") if __name__ == '__main__': # 替换为你的文件路径 go_obo_path = 'go-basic.obo' go_txt_path = 'file.txt' output_excel = 'go_level2_statistics.xlsx' go_terms = parse_obo(go_obo_path) go_freq = count_go_occurrences(go_txt_path) generate_excel(go_terms, go_freq, output_excel)
关键修复说明
- 明确层级定义:以GO三大根节点为Level 1,其直接子节点为Level 2,确保所有下属GO term都映射到对应的Level 2分类
- 完整路径追溯:遍历
is_a关系追溯到根节点,通过路径反转准确定位Level 2节点,而非仅取直接父节点 - 异常处理:对不存在的GO ID或无父节点的情况返回明确标识,避免程序崩溃
- 灵活适配:多父节点场景可调整父节点选择逻辑(当前取第一个
is_a父节点)
内容的提问来源于stack exchange,提问作者Umar
相关产品推荐
相关产品推荐

