如何通过循环遍历嵌套字典列表创建DataFrame并处理缺失键值
解决嵌套字典列表转DataFrame时的键缺失问题
问题场景
需要遍历由字典与列表嵌套组成的列表创建Pandas DataFrame,但数据中存在缺失的键值对(比如示例中第二条数据缺少indexed键),直接通过[key]访问会触发KeyError,希望缺失键时自动返回空值(Null)。
示例数据
sample_list = [{'sitemap': [{'path': 'http://test.com', 'errors': '0', 'contents': [{'type': 'web', 'submitted': '34801', 'indexed': '4656'}]}]}, {'sitemap': [{'path': 'https://example.com', 'errors': '0', 'contents': [{'type': 'web', 'submitted': '2329'}]}]}]
原尝试代码(存在问题)
import pandas as pd data_for_df = [] for each in sample_list: temp = [] temp.append(each['sitemap'][0]['path']) temp.append(each['sitemap'][0]['errors']) temp.append(each['sitemap'][0]['contents'][0]['type']) temp.append(each['sitemap'][0]['contents'][0]['submitted']) temp.append(each['sitemap'][0]['contents'][0]['indexed']) # 缺失键会触发报错 data_for_df.append(temp) df = pd.DataFrame(data_for_df, columns=['path','lastSubmitted','type','submitted'])
解决方案
核心是使用字典的get()方法安全访问键值,该方法允许指定键不存在时的默认返回值(默认是None,转DataFrame后会自动转为NaN,即空值)。
方法1:修改原循环为字典存储(推荐)
改用字典存储每行数据,列名与值一一对应,更直观且不易出错:
import pandas as pd sample_list = [{'sitemap': [{'path': 'http://test.com', 'errors': '0', 'contents': [{'type': 'web', 'submitted': '34801', 'indexed': '4656'}]}]}, {'sitemap': [{'path': 'https://example.com', 'errors': '0', 'contents': [{'type': 'web', 'submitted': '2329'}]}]}] data_for_df = [] for item in sample_list: # 提取嵌套层级的基础数据(假设sitemap和contents列表至少有一个元素) sitemap_info = item['sitemap'][0] content_info = sitemap_info['contents'][0] # 用get()安全获取值,缺失键返回None row = { 'path': sitemap_info.get('path'), 'lastSubmitted': sitemap_info.get('errors'), 'type': content_info.get('type'), 'submitted': content_info.get('submitted'), 'indexed': content_info.get('indexed') # 缺失时自动返回None } data_for_df.append(row) df = pd.DataFrame(data_for_df) print(df)
运行结果:
path lastSubmitted type submitted indexed 0 http://test.com 0 web 34801 4656.0 1 https://example.com 0 web 2329 NaN
方法2:列表推导式简化代码
如果逻辑简单,也可以用列表推导式直接生成字典列表:
import pandas as pd sample_list = [{'sitemap': [{'path': 'http://test.com', 'errors': '0', 'contents': [{'type': 'web', 'submitted': '34801', 'indexed': '4656'}]}]}, {'sitemap': [{'path': 'https://example.com', 'errors': '0', 'contents': [{'type': 'web', 'submitted': '2329'}]}]}] data_for_df = [ { 'path': item['sitemap'][0].get('path'), 'lastSubmitted': item['sitemap'][0].get('errors'), 'type': item['sitemap'][0]['contents'][0].get('type'), 'submitted': item['sitemap'][0]['contents'][0].get('submitted'), 'indexed': item['sitemap'][0]['contents'][0].get('indexed') } for item in sample_list ] df = pd.DataFrame(data_for_df)
额外提示
如果sitemap或contents列表可能为空,可以进一步用get()处理层级,避免索引报错:
sitemap_info = item.get('sitemap', [{}])[0] # 若sitemap为空,返回空字典 content_info = sitemap_info.get('contents', [{}])[0]
内容的提问来源于stack exchange,提问作者Anchobi_codes
相关产品推荐
相关产品推荐

