You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何转置列中唯一条目并关联统计基因数据(含多值分割)

处理基因-生物过程映射文件的需求与实现

需求

现有一个制表符分隔的input.txt文件,包含Gene(基因)和Biological_Process(生物过程)两列。需完成以下操作:

  • 将第二列内容按分号;分割为独立条目
  • 提取所有唯一的生物过程条目作为新的第一列
  • 为每个生物过程列出所有匹配的基因名称

输入文件(input.txt)

Gene	Biological_Process
BALF2	metabolic process
CHD4	cell organization and biogenesis;metabolic process;regulation of biological process
TCOF1	cell organization and biogenesis;regulation of biological process;transport
TOP1	cell death;cell division;cell organization and biogenesis;metabolic process;regulation of biological process;response to stimulus
BcLF1	0
BALF5	metabolic process
MTA2	cell organization and biogenesis;metabolic process;regulation of biological process
MSH6	cell organization and biogenesis;metabolic process;regulation of biological process;response to stimulus

预期输出

Biological_Process	Gene
metabolic process	BALF2	CHD4	TOP1	BALF5	MTA2	MSH6
cell organization and biogenesis	CHD4	TCOF1	TOP1	MTA2	MSH6
regulation of biological process	CHD4	TCOF1	TOP1	MTA2	MSH6
transport	TCOF1
cell death	TOP1
cell division	TOP1
response to stimulus	TOP1	MSH6

解决方案

Python 脚本实现

# 构建生物过程到基因的映射字典
process_gene_map = {}

# 读取输入文件
with open("input.txt", "r") as infile:
    # 跳过表头行
    next(infile)
    for line in infile:
        line = line.strip()
        if not line:
            continue
        # 按制表符分割基因和生物过程列
        gene, processes = line.split("\t", 1)
        # 跳过生物过程为"0"的条目
        if processes == "0":
            continue
        # 分割生物过程列表
        process_list = processes.split(";")
        for proc in process_list:
            proc = proc.strip()
            # 初始化或添加基因到对应过程
            if proc not in process_gene_map:
                process_gene_map[proc] = []
            if gene not in process_gene_map[proc]:
                process_gene_map[proc].append(gene)

# 写入输出文件
with open("output.txt", "w") as outfile:
    outfile.write("Biological_Process\tGene\n")
    # 按生物过程名称排序输出
    for proc in sorted(process_gene_map.keys()):
        gene_str = "\t".join(process_gene_map[proc])
        outfile.write(f"{proc}\t{gene_str}\n")

Awk 命令行实现

适合无需编写脚本的快速处理场景:

BEGIN {
    FS = "\t"
    OFS = "\t"
    print "Biological_Process", "Gene"
}
NR > 1 {
    if ($2 == "0") next
    n = split($2, procs, ";")
    for (i=1; i<=n; i++) {
        proc = procs[i]
        gsub(/^[[:space:]]+|[[:space:]]+$/, "", proc)
        if (!(proc in gene_list)) {
            gene_list[proc] = $1
        } else if (index("\t" gene_list[proc] "\t", "\t" $1 "\t") == 0) {
            gene_list[proc] = gene_list[proc] OFS $1
        }
    }
}
END {
    PROCINFO["sorted_in"] = "@ind_str_asc"
    for (proc in gene_list) {
        print proc, gene_list[proc]
    }
}

执行命令:

awk -f process_mapping.awk input.txt > output.txt

内容的提问来源于stack exchange,提问作者Ibk

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.12 23:30:18