You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何汇总同一GeneID的多GO Term?Python脚本输出优化求助

解决基因GO术语汇总问题

核心思路是用字典聚合每个基因对应的所有GO术语,处理完所有输入行后再统一输出,避免重复打印基因ID。

修改后的代码如下:

# 初始化字典,键为GeneID,值为对应GO术语的列表
gene_go_map = {}

for line in ref:
    if 'Protein GO term Score' in line:
        continue
    # 分割行并去除多余空白,避免格式问题
    line_parts = line.strip().split('\t')
    gene_id = line_parts[0]
    go_id = line_parts[1].strip()

    # 将GO术语添加到对应基因的列表中
    if gene_id not in gene_go_map:
        gene_go_map[gene_id] = []
    gene_go_map[gene_id].append(go_id)

# 写入汇总结果
for gene_id, go_terms in gene_go_map.items():
    # 用分号分隔多个GO术语,可根据需求替换为逗号等其他分隔符
    combined_go = '; '.join(go_terms)
    results.write(f'Znev\t{gene_id}\t{combined_go}\n')

额外优化建议

如果输入中存在重复的GO术语,可改用集合自动去重,输出时还能排序保证格式一致性:

gene_go_map = {}

for line in ref:
    if 'Protein GO term Score' in line:
        continue
    line_parts = line.strip().split('\t')
    gene_id = line_parts[0]
    go_id = line_parts[1].strip()

    if gene_id not in gene_go_map:
        gene_go_map[gene_id] = set()  # 集合自动去重
    gene_go_map[gene_id].add(go_id)

# 转成列表并排序后输出
for gene_id, go_terms in gene_go_map.items():
    combined_go = '; '.join(sorted(go_terms))
    results.write(f'Znev\t{gene_id}\t{combined_go}\n')

内容的提问来源于stack exchange,提问作者cloud1

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.20 01:45:38