JSON转CSV遇AttributeError,需修改Python脚本提取双基因列
问题描述
需要将多个指定格式的JSON文件转换为CSV,要求提取:
- 双基因列:
gene_name_A、gene_name_B - 对应摘要:
Summary_Gene_A、Summary_Gene_B - 6项评分:
well_characterized、biology、cancer、tumor、drug_targets、clinical_marker
运行原脚本时出现AttributeError,错误信息如下:
--------------------------------------------------------------------------- AttributeError Traceback (most recent call last) Cell In[4], line 76 74 print (k) 75 q1 = json.load(open(k,'r')) ---> 76 tmpDF = get_qScore(q1,mainkey_qFLU, subkey=subkey_name) 77 q1_DF = pd.concat([q1_DF,tmpDF],axis=0) 79 print (q1_DF.shape) Cell In[4], line 60, in get_qScore(q, question_dict, subkey) 58 q_scores = [] 59 for gname in q.keys(): ---> 60 for model in q[gname].keys(): 61 kx = convert_stringtodict(q[gname][model],question_dict) 62 kx.update({'gene_name':gname, 'runID':model, "model_version":model.lstrip("datacs-poc").split("_")[0], "subjectKey":subkey,}) AttributeError: 'int' object has no attribute 'keys'
示例JSON文件
{ "KAT2A": { "official_gene_symbol": "KAT2A", "brief_summary": "KAT2A, also known as GCN5, is a histone acetyltransferase that plays a key role." }, "E2F1": { "official_gene_symbol": "E2F1", "brief_summary": "E2F1 is a member of the E2F family of transcription factors." }, "The genes encode molecules that are well characterized protein-protein pairs": 7, "The protein-protein pair is relevant to biology": 8, "The protein-protein pair is relevant to cancer": 3, "The protein-protein pair is relevant to interactions between tumor": 4, "One or both of the genes are known drug targets": 5, "One or both of the genes are known clinical marker": 6 }
期望CSV格式
#> gene_name_A gene_name_B Summary_Gene_A #> 1 KAT2A E2F1 KAT2A, also known as GCN5. #> 2 KRT30 KRT31 Keratin 30 is a cytokeratin. #> Summary_Gene_B well_characterized #> 1 E2F1 is a member of the E2F family of transcription. 7 #> 2 Keratin 31 is a cytokeratin. 7 #> Biology Cancer Tumor drug_targets clinical_biomarkers #> 1 8 3 4 5 6 #> 2 8 2 3 1 5
原脚本
# test azure import sys, time, json # from openai import OpenAI import pandas as pd import re from glob import glob # define key dictionary for each question for concrete formatting mainkey_qFLU = {'summary':'Summary', 'well characterized':'well_characterized', 'biology':'Biology', 'cancer':'Cancer', 'tumor':'tumor', 'drug targets':'drug_targets', 'clinical markers':'clinical_markers' } def find_keyword(sline, keyLib): for mk in keyLib.keys(): # Regular expression pattern to find all combinations of the letters in 'gene' pattern = r'{}'.format(mk) # Finding all matches in the sample text matches = re.findall(pattern, sline, re.IGNORECASE) if matches: return keyLib[mk] else: next return False def convert_stringtodict(lines, keylib): dict_line = {} for k in lines: ksplit = k.split(":") if len(ksplit) ==2: key_tmp = find_keyword(ksplit[0].strip("\'|\"|', |").strip(), keylib) val_tmp = ksplit[1].strip("\'|\"|',|{|} ").strip() if key_tmp and val_tmp: if key_tmp == "Summary": dict_line[key_tmp] = val_tmp else: try: dict_line[key_tmp] = float(val_tmp) except: dict_line[key_tmp] = 0 else: next # print ("error in ", ksplit) return dict_line def get_qScore(q, question_dict, subkey): q_scores = [] for gname in q.keys(): for model in q[gname].keys(): kx = convert_stringtodict(q[gname][model],question_dict) kx.update({'gene_name':gname, 'runID':model, "model_version":model.lstrip("datasvc-openai-testinglab-poc-").split("_")[0], "subjectKey":subkey,}) q_scores.append(kx) print (len(q_scores)) return pd.DataFrame(q_scores) q1_DF = pd.DataFrame() for k in glob("/Users/Documents/Projects/json_files/*.json"): subkey_name = "-".join(k.split("/")[-1].split("_")[1:3]) print (k) q1 = json.load(open(k,'r')) tmpDF = get_qScore(q1,mainkey_qFLU, subkey=subkey_name) q1_DF = pd.concat([q1_DF,tmpDF],axis=0) print (q1_DF.shape) q1_DF.to_csv("./Score_parse_1_2.csv")
修改后的脚本
import json import pandas as pd from glob import glob # 定义评分项的映射关系,匹配JSON中的键到CSV列名 score_mapping = { "The genes encode molecules that are well characterized protein-protein pairs": "well_characterized", "The protein-protein pair is relevant to biology": "Biology", "The protein-protein pair is relevant to cancer": "Cancer", "The protein-protein pair is relevant to interactions between tumor": "Tumor", "One or both of the genes are known drug targets": "drug_targets", "One or both of the genes are known clinical marker": "clinical_markers" } def process_single_json(json_data): # 分离基因信息和评分信息 gene_entries = {} scores = {} for key, value in json_data.items(): if isinstance(value, dict) and "official_gene_symbol" in value and "brief_summary" in value: # 识别基因条目 gene_entries[key] = { "symbol": value["official_gene_symbol"], "summary": value["brief_summary"] } elif key in score_mapping: # 识别评分条目 scores[score_mapping[key]] = value # 提取两个基因(假设每个JSON固定两个基因) gene_list = list(gene_entries.values()) if len(gene_list) != 2: # 处理基因数量不符的情况,返回空字典 return None # 构建单行数据 row = { "gene_name_A": gene_list[0]["symbol"], "Summary_Gene_A": gene_list[0]["summary"], "gene_name_B": gene_list[1]["symbol"], "Summary_Gene_B": gene_list[1]["summary"] } # 合并评分数据 row.update(scores) return row # 主处理流程 q1_DF = pd.DataFrame() json_files = glob("/Users/Documents/Projects/json_files/*.json") for file_path in json_files: print(f"Processing file: {file_path}") with open(file_path, 'r') as f: json_data = json.load(f) row_data = process_single_json(json_data) if row_data: tmpDF = pd.DataFrame([row_data]) q1_DF = pd.concat([q1_DF, tmpDF], axis=0) # 重置索引 q1_DF = q1_DF.reset_index(drop=True) print(f"Final DataFrame shape: {q1_DF.shape}") # 保存为CSV q1_DF.to_csv("./Score_parse_1_2.csv", index=False)
修改说明
- 解决报错问题:原脚本遍历JSON所有键并尝试调用
.keys(),但评分项的值是整数,不是字典。修改后先区分基因条目(字典类型且包含指定键)和评分条目,避免对非字典值调用字典方法。 - 实现双基因提取:从JSON中筛选出基因条目,固定提取前两个基因(假设每个JSON对应一对基因),分别填充
gene_name_A、Summary_Gene_A和gene_name_B、Summary_Gene_B列。 - 简化评分提取:直接通过预定义的
score_mapping匹配JSON中的评分键到目标列名,无需复杂的正则匹配,提升效率和准确性。 - 增加容错处理:如果JSON中的基因数量不是2个,跳过该文件的处理,避免程序崩溃。
内容的提问来源于stack exchange,提问作者Abdul Khayum
相关产品推荐
相关产品推荐

