如何从内层为列表的嵌套字典构建指定结构的pandas DataFrame
实现代码
import pandas as pd row_list = [] # 遍历外层字典的所有基因ID及对应子条目列表 for gene_id, subunit_entries in annot_dict.items(): # 遍历每个转录本/基因子字典 for sub_item in subunit_entries: # 取出子字典唯一的键(subunit ID)和对应属性列表 subunit_id, attr_list = next(iter(sub_item.items())) # 按属性顺序解包:biotype、start index、end index、strand、description biotype, start_idx, end_idx, strand, desc = attr_list # 组装行数据 row_list.append({ "subunit_ID": subunit_id, "gene_ID": gene_id, "start_index": start_idx, "end_index": end_idx, "strand": strand, "biotype": biotype, "desc": desc }) # 转换为DataFrame df = pd.DataFrame(row_list) # 若需严格匹配你要求的列顺序,可补充以下代码 df = df[["subunit_ID", "gene_ID", "start_index", "end_index", "strand", "biotype", "desc"]]
注意事项
属性解包顺序和你描述的[biotype、start index、end index、strand、description]对齐,若你的gff3解析函数输出的属性列表顺序有调整,直接修改解包变量的顺序即可。
内容的提问来源于stack exchange,提问作者Logan
相关产品推荐
相关产品推荐

