如何转置列中唯一条目并关联统计基因数据(含多值分割)
处理基因-生物过程映射文件的需求与实现
需求
现有一个制表符分隔的input.txt文件,包含Gene(基因)和Biological_Process(生物过程)两列。需完成以下操作:
- 将第二列内容按分号
;分割为独立条目 - 提取所有唯一的生物过程条目作为新的第一列
- 为每个生物过程列出所有匹配的基因名称
输入文件(input.txt)
Gene Biological_Process BALF2 metabolic process CHD4 cell organization and biogenesis;metabolic process;regulation of biological process TCOF1 cell organization and biogenesis;regulation of biological process;transport TOP1 cell death;cell division;cell organization and biogenesis;metabolic process;regulation of biological process;response to stimulus BcLF1 0 BALF5 metabolic process MTA2 cell organization and biogenesis;metabolic process;regulation of biological process MSH6 cell organization and biogenesis;metabolic process;regulation of biological process;response to stimulus
预期输出
Biological_Process Gene metabolic process BALF2 CHD4 TOP1 BALF5 MTA2 MSH6 cell organization and biogenesis CHD4 TCOF1 TOP1 MTA2 MSH6 regulation of biological process CHD4 TCOF1 TOP1 MTA2 MSH6 transport TCOF1 cell death TOP1 cell division TOP1 response to stimulus TOP1 MSH6
解决方案
Python 脚本实现
# 构建生物过程到基因的映射字典 process_gene_map = {} # 读取输入文件 with open("input.txt", "r") as infile: # 跳过表头行 next(infile) for line in infile: line = line.strip() if not line: continue # 按制表符分割基因和生物过程列 gene, processes = line.split("\t", 1) # 跳过生物过程为"0"的条目 if processes == "0": continue # 分割生物过程列表 process_list = processes.split(";") for proc in process_list: proc = proc.strip() # 初始化或添加基因到对应过程 if proc not in process_gene_map: process_gene_map[proc] = [] if gene not in process_gene_map[proc]: process_gene_map[proc].append(gene) # 写入输出文件 with open("output.txt", "w") as outfile: outfile.write("Biological_Process\tGene\n") # 按生物过程名称排序输出 for proc in sorted(process_gene_map.keys()): gene_str = "\t".join(process_gene_map[proc]) outfile.write(f"{proc}\t{gene_str}\n")
Awk 命令行实现
适合无需编写脚本的快速处理场景:
BEGIN { FS = "\t" OFS = "\t" print "Biological_Process", "Gene" } NR > 1 { if ($2 == "0") next n = split($2, procs, ";") for (i=1; i<=n; i++) { proc = procs[i] gsub(/^[[:space:]]+|[[:space:]]+$/, "", proc) if (!(proc in gene_list)) { gene_list[proc] = $1 } else if (index("\t" gene_list[proc] "\t", "\t" $1 "\t") == 0) { gene_list[proc] = gene_list[proc] OFS $1 } } } END { PROCINFO["sorted_in"] = "@ind_str_asc" for (proc in gene_list) { print proc, gene_list[proc] } }
执行命令:
awk -f process_mapping.awk input.txt > output.txt
内容的提问来源于stack exchange,提问作者Ibk
相关产品推荐
相关产品推荐

