如何在嵌套字典中转换protein accession列表为字符串及数值列表为整数?
解决方案
可以写个简单的Python函数批量处理这个嵌套字典,逻辑直接明了:遍历每个蛋白的详情字典,针对不同字段做对应转换。
代码实现
# 示例输入字典 protein_data = { "A0A0H3LJT0_e": { "protein accession": ["A0A0H3LJT0_e"], "sequence length": ["102"], "analysis": ["SMART"], "signature accession": ["SM00886"], "signature description": ["Dabb_2"], "start location": ["4"], "stop location": ["98"], "e-value": ["1.5E-22"], "interpro accession": ["IPR013097"], "interpro description": ["Stress responsive alpha-beta barrel"], "nunique": [2], "domain_count": [1], } } def process_protein_dict(input_dict): # 定义需要转成整数列表的字段 int_convert_fields = ['sequence length', 'start location', 'stop location', 'nunique', 'domain_count'] for protein_id, details in input_dict.items(): # 处理protein accession:单元素列表转字符串 if 'protein accession' in details: details['protein accession'] = details['protein accession'][0] # 处理需要转整数的字段:字符串列表转整数列表 for field in int_convert_fields: if field in details: details[field] = [int(item) for item in details[field]] return input_dict # 调用函数处理 processed_data = process_protein_dict(protein_data) # 打印结果查看 import pprint pprint.pprint(processed_data)
输出结果
{ 'A0A0H3LJT0_e': { 'analysis': ['SMART'], 'domain_count': [1], 'e-value': ['1.5E-22'], 'interpro accession': ['IPR013097'], 'interpro description': ['Stress responsive alpha-beta barrel'], 'nunique': [2], 'protein accession': 'A0A0H3LJT0_e', 'sequence length': [102], 'signature accession': ['SM00886'], 'signature description': ['Dabb_2'], 'start location': [4], 'stop location': [98] } }
说明
- 要是后续需要新增转整数的字段,直接往
int_convert_fields列表里加就行 - 代码做了字段存在性判断,避免字典缺字段时报错
- 就算字段对应的列表有多元素,这个逻辑也适用(比如
sequence length是['102','200'],会转成[102,200])
内容的提问来源于stack exchange,提问作者Aurinko
相关产品推荐
相关产品推荐

