You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何清理嵌套字典中的重复值并处理None值?

处理嵌套字典中的重复值与None值

针对你给出的嵌套字典结构,以下是两种实用的处理方案,分别对应不同的数据使用需求:

方案1:简化重复列表为单一值,过滤None保留有效条目

对于全重复的列表(如protein accession、sequence length)提取唯一值;对于混有None的列表(如start location、e-value)过滤掉None,保留有效数据,同时保留原字典的层级结构:

def clean_protein_dict(protein_data):
    cleaned = {}
    for protein_id, details in protein_data.items():
        cleaned_details = {}
        for key, values in details.items():
            # 过滤列表中的None值
            filtered_vals = [v for v in values if v is not None]
            # 若过滤后所有值一致,转为单一值;否则保留过滤后的列表
            if len(set(filtered_vals)) == 1:
                cleaned_details[key] = filtered_vals[0]
            else:
                cleaned_details[key] = filtered_vals
        cleaned[protein_id] = cleaned_details
    return cleaned

# 测试示例
original_dict = {
    'C4QY10_e': {
        'protein accession': ['C4QY10_e', 'C4QY10_e', 'C4QY10_e', 'C4QY10_e', 'C4QY10_e'],
        'sequence length': ['1879', '1879', '1879', '1879', '1879'],
        'analysis': ['Pfam', 'Pfam', 'Pfam', 'Pfam', 'Pfam'],
        'signature accession': ['PF18314', 'PF02801', 'PF18325', 'PF00109', 'PF01648'],
        'signature description': ['Fatty acid synthase type I helical domain', 'Beta-ketoacyl synthase', 'Fatty acid synthase subunit alpha Acyl carrier domain', 'Beta-ketoacyl synthase', "4'-phosphopantetheinyl transferase superfamily"],
        'start location': ['328', None, '139', None, '1761'],
        'stop location': ['528', None, '300', None, '1861'],
        'e-value': ['4.7E-73', None, '1.3E-72', None, '1.4E-18'],
        'interpro accession': ['IPR041550', None, 'IPR040899', None, 'IPR008278'],
        'interpro description': ['Fatty acid synthase type I', None, 'Fatty acid synthase subunit alpha', None, "4'-phosphopantetheinyl transferase domain"],
        'nunique': [1, 1, 1, 1, 1],
        'domain_count': [5, 5, 5, 5, 5]
    }
}

cleaned_result = clean_protein_dict(original_dict)
print(cleaned_result['C4QY10_e'])

处理后,protein accession会转为字符串'C4QY10_e',start location会变成['328', '139', '1761'],既去除冗余又保留有效信息。

方案2:重构数据结构,按结构域分组

如果你的数据本质是一个蛋白对应多个结构域,原列表是按结构域位置对齐的,那么将每个结构域的信息单独封装为字典会更便于后续处理:

def restructure_protein_data(protein_data):
    restructured = {}
    for protein_id, details in protein_data.items():
        domain_count = details['domain_count'][0]
        domains = []
        # 遍历每个结构域的位置索引
        for idx in range(domain_count):
            domain_info = {}
            for key, values in details.items():
                val = values[idx]
                # 保留非None字段,如需保留None可删除此判断
                if val is not None:
                    domain_info[key] = val
            domains.append(domain_info)
        restructured[protein_id] = {
            'protein accession': details['protein accession'][0],
            'sequence length': details['sequence length'][0],
            'analysis': details['analysis'][0],
            'domains': domains
        }
    return restructured

# 测试示例
restructured_result = restructure_protein_data(original_dict)
print(restructured_result['C4QY10_e']['domains'][0])

处理后,每个结构域的信息会独立成字典,比如第一个结构域的信息如下:

{
    'protein accession': 'C4QY10_e',
    'sequence length': '1879',
    'analysis': 'Pfam',
    'signature accession': 'PF18314',
    'signature description': 'Fatty acid synthase type I helical domain',
    'start location': '328',
    'stop location': '528',
    'e-value': '4.7E-73',
    'interpro accession': 'IPR041550',
    'interpro description': 'Fatty acid synthase type I',
    'nunique': 1,
    'domain_count': 5
}

额外提示

  • 若需要将字符串类型的数值(如sequence length)转为整数,可在处理时添加int(filtered_vals[0])转换
  • 若需保留None对应的字段(比如标记结构域缺失的信息),可删除代码中过滤None的逻辑,仅处理全重复列表

内容的提问来源于stack exchange,提问作者Aurinko

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.16 22:30:48