You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

JSON转CSV遇AttributeError,需修改Python脚本提取双基因列

问题描述

需要将多个指定格式的JSON文件转换为CSV,要求提取:

  • 双基因列:gene_name_A、gene_name_B
  • 对应摘要:Summary_Gene_A、Summary_Gene_B
  • 6项评分:well_characterized、biology、cancer、tumor、drug_targets、clinical_marker

运行原脚本时出现AttributeError,错误信息如下:

---------------------------------------------------------------------------
AttributeError                            Traceback (most recent call last)
Cell In[4], line 76
     74     print (k)
     75     q1 = json.load(open(k,'r'))
---> 76     tmpDF = get_qScore(q1,mainkey_qFLU, subkey=subkey_name)
     77     q1_DF = pd.concat([q1_DF,tmpDF],axis=0)
     79 print (q1_DF.shape)

Cell In[4], line 60, in get_qScore(q, question_dict, subkey)
     58 q_scores = []
     59 for gname in q.keys():
---> 60     for model in q[gname].keys():
     61         kx = convert_stringtodict(q[gname][model],question_dict)
     62         kx.update({'gene_name':gname,
                     'runID':model,
                     "model_version":model.lstrip("datacs-poc").split("_")[0],
                     "subjectKey":subkey,})

AttributeError: 'int' object has no attribute 'keys'

示例JSON文件

{
    "KAT2A": {
        "official_gene_symbol": "KAT2A",
        "brief_summary": "KAT2A, also known as GCN5, is a histone acetyltransferase that plays a key role."
    },
    "E2F1": {
        "official_gene_symbol": "E2F1",
        "brief_summary": "E2F1 is a member of the E2F family of transcription factors."
    },
    "The genes encode molecules that are well characterized protein-protein pairs": 7,
    "The protein-protein pair is relevant to biology": 8,
    "The protein-protein pair is relevant to cancer": 3,
    "The protein-protein pair is relevant to interactions between tumor": 4,
    "One or both of the genes are known drug targets": 5,
    "One or both of the genes are known clinical marker": 6
}

期望CSV格式

#>   gene_name_A gene_name_B               Summary_Gene_A
#> 1       KAT2A        E2F1   KAT2A, also known as GCN5.
#> 2       KRT30       KRT31 Keratin 30 is a cytokeratin.
#>                                         Summary_Gene_B well_characterized
#> 1 E2F1 is a member of the E2F family of transcription.                  7
#> 2                         Keratin 31 is a cytokeratin.                  7
#>   Biology Cancer Tumor drug_targets clinical_biomarkers
#> 1       8      3     4            5                   6
#> 2       8      2     3            1                   5

原脚本

# test azure
import sys, time, json
# from openai import OpenAI
import pandas as pd
import re
from glob import glob
# define key dictionary for each question for concrete formatting

mainkey_qFLU = {'summary':'Summary',
                'well characterized':'well_characterized',
                'biology':'Biology',
                'cancer':'Cancer',
                'tumor':'tumor',
                'drug targets':'drug_targets',
                'clinical markers':'clinical_markers'
               }


def find_keyword(sline, keyLib):
    for mk in keyLib.keys():

        # Regular expression pattern to find all combinations of the letters in 'gene'
        pattern = r'{}'.format(mk)

        # Finding all matches in the sample text
        matches = re.findall(pattern, sline, re.IGNORECASE)
        if matches:
            return keyLib[mk]
        else:
            next
    return False



def convert_stringtodict(lines, keylib):
    dict_line = {}
    for k in lines:
        ksplit = k.split(":")

        if len(ksplit) ==2:
            key_tmp = find_keyword(ksplit[0].strip("\'|\"|', |").strip(), keylib)
            val_tmp = ksplit[1].strip("\'|\"|',|{|} ").strip()
            if key_tmp and val_tmp:
                if key_tmp == "Summary":
                    dict_line[key_tmp] = val_tmp
                else:
                    try:
                        dict_line[key_tmp] = float(val_tmp)
                    except:
                        dict_line[key_tmp] = 0
            else:
                next
                # print ("error in ", ksplit)

    return dict_line

def get_qScore(q, question_dict, subkey):
    q_scores = []
    for gname in q.keys():
        for model in q[gname].keys():
            kx = convert_stringtodict(q[gname][model],question_dict)
            kx.update({'gene_name':gname,
                        'runID':model,
                        "model_version":model.lstrip("datasvc-openai-testinglab-poc-").split("_")[0],
                        "subjectKey":subkey,})
            q_scores.append(kx)
    print (len(q_scores))
    return pd.DataFrame(q_scores)


q1_DF = pd.DataFrame()
for k in glob("/Users/Documents/Projects/json_files/*.json"):
    subkey_name = "-".join(k.split("/")[-1].split("_")[1:3])
    print (k)
    q1 = json.load(open(k,'r'))
    tmpDF = get_qScore(q1,mainkey_qFLU, subkey=subkey_name)
    q1_DF = pd.concat([q1_DF,tmpDF],axis=0)

print (q1_DF.shape)

q1_DF.to_csv("./Score_parse_1_2.csv")
修改后的脚本
import json
import pandas as pd
from glob import glob

# 定义评分项的映射关系,匹配JSON中的键到CSV列名
score_mapping = {
    "The genes encode molecules that are well characterized protein-protein pairs": "well_characterized",
    "The protein-protein pair is relevant to biology": "Biology",
    "The protein-protein pair is relevant to cancer": "Cancer",
    "The protein-protein pair is relevant to interactions between tumor": "Tumor",
    "One or both of the genes are known drug targets": "drug_targets",
    "One or both of the genes are known clinical marker": "clinical_markers"
}

def process_single_json(json_data):
    # 分离基因信息和评分信息
    gene_entries = {}
    scores = {}
    
    for key, value in json_data.items():
        if isinstance(value, dict) and "official_gene_symbol" in value and "brief_summary" in value:
            # 识别基因条目
            gene_entries[key] = {
                "symbol": value["official_gene_symbol"],
                "summary": value["brief_summary"]
            }
        elif key in score_mapping:
            # 识别评分条目
            scores[score_mapping[key]] = value
    
    # 提取两个基因(假设每个JSON固定两个基因)
    gene_list = list(gene_entries.values())
    if len(gene_list) != 2:
        # 处理基因数量不符的情况,返回空字典
        return None
    
    # 构建单行数据
    row = {
        "gene_name_A": gene_list[0]["symbol"],
        "Summary_Gene_A": gene_list[0]["summary"],
        "gene_name_B": gene_list[1]["symbol"],
        "Summary_Gene_B": gene_list[1]["summary"]
    }
    # 合并评分数据
    row.update(scores)
    return row

# 主处理流程
q1_DF = pd.DataFrame()
json_files = glob("/Users/Documents/Projects/json_files/*.json")

for file_path in json_files:
    print(f"Processing file: {file_path}")
    with open(file_path, 'r') as f:
        json_data = json.load(f)
    
    row_data = process_single_json(json_data)
    if row_data:
        tmpDF = pd.DataFrame([row_data])
        q1_DF = pd.concat([q1_DF, tmpDF], axis=0)

# 重置索引
q1_DF = q1_DF.reset_index(drop=True)
print(f"Final DataFrame shape: {q1_DF.shape}")

# 保存为CSV
q1_DF.to_csv("./Score_parse_1_2.csv", index=False)
修改说明
  1. 解决报错问题:原脚本遍历JSON所有键并尝试调用.keys(),但评分项的值是整数,不是字典。修改后先区分基因条目(字典类型且包含指定键)和评分条目,避免对非字典值调用字典方法。
  2. 实现双基因提取:从JSON中筛选出基因条目,固定提取前两个基因(假设每个JSON对应一对基因),分别填充gene_name_A、Summary_Gene_A和gene_name_B、Summary_Gene_B列。
  3. 简化评分提取:直接通过预定义的score_mapping匹配JSON中的评分键到目标列名,无需复杂的正则匹配,提升效率和准确性。
  4. 增加容错处理:如果JSON中的基因数量不是2个,跳过该文件的处理,避免程序崩溃。

内容的提问来源于stack exchange,提问作者Abdul Khayum

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.30 23:02:02