You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何匹配JSON与CSV中的基因位点并关联抗生素与文件名?

解决CSV与JSON文件的基因位点匹配问题

需求概述

需要处理两类文件,实现基因位点的匹配并输出关联信息:

  • 多个空格分隔的CSV/TSV文件:每行包含Gene_Aminoacids(基因位点突变)和Filename(样本文件名),格式示例:
Gene_Aminoacids Filename
gyrA_S95T   SRR9851427
tlyA_L11L   SRR9851427
katG_R463L  SRR9851427
  • JSON文件:键为基因位点突变,值为对应的抗生素列表,示例片段:
{
  "gyrA_A74S": ["Quinolones"],
  "gyrA_D89X": ["Quinolones"],
  "tlyA_C-83T": ["Capreomycin"],
  "katG_R104Q": ["Isoniazid"],
  "katG_S315I": ["Isoniazid"]
}

最终目标是输出包含基因位点、对应抗生素、样本文件名的匹配结果,示例格式:

Gene_Aminoacids | Antibiotic | Filename
katG_R104Q      | Isoniazid  | SRR9851427

现有代码的问题

你的代码存在几个逻辑缺陷:

  • glob.glob()返回的是文件路径列表,samp作为列表直接拼接路径会报错
  • 仅读取JSON的键,未保留抗生素关联信息,无法完成后续匹配
  • 文件读取仅打印内容,未提取并存储基因位点与文件名的对应关系

完整解决方案

步骤1:加载抗生素映射字典

保留JSON完整数据,方便后续通过基因位点直接查询对应抗生素:

import json
import glob
import csv

def load_antibiotic_map(json_file):
    with open(json_file, 'r') as f:
        return json.load(f)

# 加载JSON中的抗生素映射关系
antibiotic_map = load_antibiotic_map("tb_TEST.json")

步骤2:批量获取目标文件路径

用glob直接匹配所有目标txt文件,避免路径拼接错误:

# 设置目标文件的路径匹配模式
file_pattern = "Replaced_P_G.ann.vcf/*_G.P.vcf_replaced.txt"
# 获取所有符合条件的文件完整路径
target_files = glob.glob(file_pattern)

步骤3:提取数据并完成匹配

遍历每个文件,读取每行的基因位点和样本名,与JSON中的数据匹配:

results = []

for file_path in target_files:
    with open(file_path, 'r') as f:
        next(f)  # 跳过表头行
        for line in f:
            # 按空格/制表符分割每行数据
            parts = line.strip().split()
            if len(parts) >= 2:
                gene_mutation = parts[0]
                sample_name = parts[1]
                # 检查当前基因位点是否在抗生素映射中
                if gene_mutation in antibiotic_map:
                    # 将抗生素列表转为字符串,方便输出
                    antibiotics = ", ".join(antibiotic_map[gene_mutation])
                    results.append((gene_mutation, antibiotics, sample_name))

步骤4:输出结果

支持控制台打印和写入CSV文件两种方式:

# 打印结果到控制台
print("Gene_Aminoacids\tAntibiotic\tFilename")
for item in results:
    print(f"{item[0]}\t{item[1]}\t{item[2]}")

# 写入结果到CSV文件(可选)
with open("mutation_antibiotic_result.csv", 'w', newline='') as csvfile:
    writer = csv.writer(csvfile, delimiter='\t')
    writer.writerow(["Gene_Aminoacids", "Antibiotic", "Filename"])
    writer.writerows(results)

完整整合代码

import json
import glob
import csv

def load_antibiotic_map(json_file):
    with open(json_file, 'r') as f:
        return json.load(f)

def main():
    # 加载抗生素映射数据
    antibiotic_map = load_antibiotic_map("tb_TEST.json")
    # 获取所有目标文件路径
    file_pattern = "Replaced_P_G.ann.vcf/*_G.P.vcf_replaced.txt"
    target_files = glob.glob(file_pattern)
    
    results = []
    
    # 遍历处理每个文件
    for file_path in target_files:
        with open(file_path, 'r') as f:
            next(f)  # 跳过表头
            for line in f:
                parts = line.strip().split()
                if len(parts) >= 2:
                    gene_mutation = parts[0]
                    sample_name = parts[1]
                    # 匹配成功则记录结果
                    if gene_mutation in antibiotic_map:
                        antibiotics = ", ".join(antibiotic_map[gene_mutation])
                        results.append((gene_mutation, antibiotics, sample_name))
    
    # 控制台输出
    print("Gene_Aminoacids\tAntibiotic\tFilename")
    for item in results:
        print(f"{item[0]}\t{item[1]}\t{item[2]}")
    
    # 写入结果文件
    with open("mutation_antibiotic_result.csv", 'w', newline='') as csvfile:
        writer = csv.writer(csvfile, delimiter='\t')
        writer.writerow(["Gene_Aminoacids", "Antibiotic", "Filename"])
        writer.writerows(results)

if __name__ == "__main__":
    main()

关键说明

  • 用split()分割行数据:兼容空格或制表符分隔的文件格式
  • 保留JSON完整字典:直接通过键查询抗生素信息,逻辑更简洁
  • 结果存储为列表:方便后续灵活处理(打印、写入文件等)
  • 跳过表头:避免将标题行误处理为数据

内容的提问来源于stack exchange,提问作者Pocket_sunshine

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.13 18:10:27