嵌套JSON转列表再转DataFrame:键检查与多键拆分问题
处理Google Books API嵌套JSON数据的解决方案
1. 在列表推导式中检查指定键是否存在
可以通过两种方式规避键缺失导致的错误:
- 先过滤再提取:在列表推导式的条件中直接判断键是否存在,过滤掉不包含目标键的条目后再提取内容:
books_data = [ { 'title': item['volumeInfo']['title'], 'industry_ids': item['volumeInfo']['industryIdentifiers'] # 按需添加其他字段 } for item in api_response.get('items', []) if 'volumeInfo' in item and 'industryIdentifiers' in item['volumeInfo'] ]
- 使用
get()方法容错:如果不想过滤条目,而是用默认值填充缺失内容,可借助字典的get()方法指定默认值(如空字典或空字符串),避免抛出KeyError:
books_data = [ { 'title': item.get('volumeInfo', {}).get('title', ''), 'industry_ids': item.get('volumeInfo', {}).get('industryIdentifiers', []) } for item in api_response.get('items', []) ]
item.get('volumeInfo', {})表示若item中无volumeInfo键,就返回空字典,后续的.get()操作不会报错。
2. 提取顺序不固定的ISBN_10和ISBN_13
针对industryIdentifiers中ISBN顺序不固定的问题,有两种简洁处理方式:
方式1:生成器表达式+next()直接提取
在列表推导式中通过生成器匹配type字段,找不到对应类型时返回空字符串:
books_data = [ { 'title': item.get('volumeInfo', {}).get('title', ''), 'isbn10': next((id['identifier'] for id in item.get('volumeInfo', {}).get('industryIdentifiers', []) if id.get('type') == 'ISBN_10'), ''), 'isbn13': next((id['identifier'] for id in item.get('volumeInfo', {}).get('industryIdentifiers', []) if id.get('type') == 'ISBN_13'), '') } for item in api_response.get('items', []) ]
方式2:定义辅助函数复用提取逻辑
如果需要多次复用ISBN提取逻辑,可单独定义函数处理,再在列表推导式中调用:
def get_isbns(industry_ids): isbn_map = {'ISBN_10': '', 'ISBN_13': ''} for id_entry in industry_ids: if id_entry.get('type') in isbn_map: isbn_map[id_entry['type']] = id_entry.get('identifier', '') return isbn_map books_data = [ { 'title': item.get('volumeInfo', {}).get('title', ''), **get_isbns(item.get('volumeInfo', {}).get('industryIdentifiers', [])) } for item in api_response.get('items', []) ]
用**语法将函数返回的字典展开到结果字典中,简化字段赋值。
内容的提问来源于stack exchange,提问作者Harold Meneley
相关产品推荐
相关产品推荐

