如何清理嵌套字典中的重复值并处理None值?
处理嵌套字典中的重复值与None值
针对你给出的嵌套字典结构,以下是两种实用的处理方案,分别对应不同的数据使用需求:
方案1:简化重复列表为单一值,过滤None保留有效条目
对于全重复的列表(如protein accession、sequence length)提取唯一值;对于混有None的列表(如start location、e-value)过滤掉None,保留有效数据,同时保留原字典的层级结构:
def clean_protein_dict(protein_data): cleaned = {} for protein_id, details in protein_data.items(): cleaned_details = {} for key, values in details.items(): # 过滤列表中的None值 filtered_vals = [v for v in values if v is not None] # 若过滤后所有值一致,转为单一值;否则保留过滤后的列表 if len(set(filtered_vals)) == 1: cleaned_details[key] = filtered_vals[0] else: cleaned_details[key] = filtered_vals cleaned[protein_id] = cleaned_details return cleaned # 测试示例 original_dict = { 'C4QY10_e': { 'protein accession': ['C4QY10_e', 'C4QY10_e', 'C4QY10_e', 'C4QY10_e', 'C4QY10_e'], 'sequence length': ['1879', '1879', '1879', '1879', '1879'], 'analysis': ['Pfam', 'Pfam', 'Pfam', 'Pfam', 'Pfam'], 'signature accession': ['PF18314', 'PF02801', 'PF18325', 'PF00109', 'PF01648'], 'signature description': ['Fatty acid synthase type I helical domain', 'Beta-ketoacyl synthase', 'Fatty acid synthase subunit alpha Acyl carrier domain', 'Beta-ketoacyl synthase', "4'-phosphopantetheinyl transferase superfamily"], 'start location': ['328', None, '139', None, '1761'], 'stop location': ['528', None, '300', None, '1861'], 'e-value': ['4.7E-73', None, '1.3E-72', None, '1.4E-18'], 'interpro accession': ['IPR041550', None, 'IPR040899', None, 'IPR008278'], 'interpro description': ['Fatty acid synthase type I', None, 'Fatty acid synthase subunit alpha', None, "4'-phosphopantetheinyl transferase domain"], 'nunique': [1, 1, 1, 1, 1], 'domain_count': [5, 5, 5, 5, 5] } } cleaned_result = clean_protein_dict(original_dict) print(cleaned_result['C4QY10_e'])
处理后,protein accession会转为字符串'C4QY10_e',start location会变成['328', '139', '1761'],既去除冗余又保留有效信息。
方案2:重构数据结构,按结构域分组
如果你的数据本质是一个蛋白对应多个结构域,原列表是按结构域位置对齐的,那么将每个结构域的信息单独封装为字典会更便于后续处理:
def restructure_protein_data(protein_data): restructured = {} for protein_id, details in protein_data.items(): domain_count = details['domain_count'][0] domains = [] # 遍历每个结构域的位置索引 for idx in range(domain_count): domain_info = {} for key, values in details.items(): val = values[idx] # 保留非None字段,如需保留None可删除此判断 if val is not None: domain_info[key] = val domains.append(domain_info) restructured[protein_id] = { 'protein accession': details['protein accession'][0], 'sequence length': details['sequence length'][0], 'analysis': details['analysis'][0], 'domains': domains } return restructured # 测试示例 restructured_result = restructure_protein_data(original_dict) print(restructured_result['C4QY10_e']['domains'][0])
处理后,每个结构域的信息会独立成字典,比如第一个结构域的信息如下:
{ 'protein accession': 'C4QY10_e', 'sequence length': '1879', 'analysis': 'Pfam', 'signature accession': 'PF18314', 'signature description': 'Fatty acid synthase type I helical domain', 'start location': '328', 'stop location': '528', 'e-value': '4.7E-73', 'interpro accession': 'IPR041550', 'interpro description': 'Fatty acid synthase type I', 'nunique': 1, 'domain_count': 5 }
额外提示
- 若需要将字符串类型的数值(如
sequence length)转为整数,可在处理时添加int(filtered_vals[0])转换 - 若需保留None对应的字段(比如标记结构域缺失的信息),可删除代码中过滤None的逻辑,仅处理全重复列表
内容的提问来源于stack exchange,提问作者Aurinko
相关产品推荐
相关产品推荐

