如何匹配JSON与CSV中的基因位点并关联抗生素与文件名?
解决CSV与JSON文件的基因位点匹配问题
需求概述
需要处理两类文件,实现基因位点的匹配并输出关联信息:
- 多个空格分隔的CSV/TSV文件:每行包含
Gene_Aminoacids(基因位点突变)和Filename(样本文件名),格式示例:
Gene_Aminoacids Filename gyrA_S95T SRR9851427 tlyA_L11L SRR9851427 katG_R463L SRR9851427
- JSON文件:键为基因位点突变,值为对应的抗生素列表,示例片段:
{ "gyrA_A74S": ["Quinolones"], "gyrA_D89X": ["Quinolones"], "tlyA_C-83T": ["Capreomycin"], "katG_R104Q": ["Isoniazid"], "katG_S315I": ["Isoniazid"] }
最终目标是输出包含基因位点、对应抗生素、样本文件名的匹配结果,示例格式:
Gene_Aminoacids | Antibiotic | Filename katG_R104Q | Isoniazid | SRR9851427
现有代码的问题
你的代码存在几个逻辑缺陷:
glob.glob()返回的是文件路径列表,samp作为列表直接拼接路径会报错- 仅读取JSON的键,未保留抗生素关联信息,无法完成后续匹配
- 文件读取仅打印内容,未提取并存储基因位点与文件名的对应关系
完整解决方案
步骤1:加载抗生素映射字典
保留JSON完整数据,方便后续通过基因位点直接查询对应抗生素:
import json import glob import csv def load_antibiotic_map(json_file): with open(json_file, 'r') as f: return json.load(f) # 加载JSON中的抗生素映射关系 antibiotic_map = load_antibiotic_map("tb_TEST.json")
步骤2:批量获取目标文件路径
用glob直接匹配所有目标txt文件,避免路径拼接错误:
# 设置目标文件的路径匹配模式 file_pattern = "Replaced_P_G.ann.vcf/*_G.P.vcf_replaced.txt" # 获取所有符合条件的文件完整路径 target_files = glob.glob(file_pattern)
步骤3:提取数据并完成匹配
遍历每个文件,读取每行的基因位点和样本名,与JSON中的数据匹配:
results = [] for file_path in target_files: with open(file_path, 'r') as f: next(f) # 跳过表头行 for line in f: # 按空格/制表符分割每行数据 parts = line.strip().split() if len(parts) >= 2: gene_mutation = parts[0] sample_name = parts[1] # 检查当前基因位点是否在抗生素映射中 if gene_mutation in antibiotic_map: # 将抗生素列表转为字符串,方便输出 antibiotics = ", ".join(antibiotic_map[gene_mutation]) results.append((gene_mutation, antibiotics, sample_name))
步骤4:输出结果
支持控制台打印和写入CSV文件两种方式:
# 打印结果到控制台 print("Gene_Aminoacids\tAntibiotic\tFilename") for item in results: print(f"{item[0]}\t{item[1]}\t{item[2]}") # 写入结果到CSV文件(可选) with open("mutation_antibiotic_result.csv", 'w', newline='') as csvfile: writer = csv.writer(csvfile, delimiter='\t') writer.writerow(["Gene_Aminoacids", "Antibiotic", "Filename"]) writer.writerows(results)
完整整合代码
import json import glob import csv def load_antibiotic_map(json_file): with open(json_file, 'r') as f: return json.load(f) def main(): # 加载抗生素映射数据 antibiotic_map = load_antibiotic_map("tb_TEST.json") # 获取所有目标文件路径 file_pattern = "Replaced_P_G.ann.vcf/*_G.P.vcf_replaced.txt" target_files = glob.glob(file_pattern) results = [] # 遍历处理每个文件 for file_path in target_files: with open(file_path, 'r') as f: next(f) # 跳过表头 for line in f: parts = line.strip().split() if len(parts) >= 2: gene_mutation = parts[0] sample_name = parts[1] # 匹配成功则记录结果 if gene_mutation in antibiotic_map: antibiotics = ", ".join(antibiotic_map[gene_mutation]) results.append((gene_mutation, antibiotics, sample_name)) # 控制台输出 print("Gene_Aminoacids\tAntibiotic\tFilename") for item in results: print(f"{item[0]}\t{item[1]}\t{item[2]}") # 写入结果文件 with open("mutation_antibiotic_result.csv", 'w', newline='') as csvfile: writer = csv.writer(csvfile, delimiter='\t') writer.writerow(["Gene_Aminoacids", "Antibiotic", "Filename"]) writer.writerows(results) if __name__ == "__main__": main()
关键说明
- 用
split()分割行数据:兼容空格或制表符分隔的文件格式 - 保留JSON完整字典:直接通过键查询抗生素信息,逻辑更简洁
- 结果存储为列表:方便后续灵活处理(打印、写入文件等)
- 跳过表头:避免将标题行误处理为数据
内容的提问来源于stack exchange,提问作者Pocket_sunshine
相关产品推荐
相关产品推荐

