嵌套字典元素提取问题:Python代码修复与键缺失处理
Fixing KeyError and Empty List Issues in Your Python Data Processing Code
Let's break down the issues in your existing code and fix them step by step:
Key Issues Identified
- Incorrect Key Case for Business: You used lowercase
businessinstead of the correct capitalizedBusinesswhen accessing the list. - Wrong Path for Policies: Policies are nested inside
_source, but your code was looking for them directly in the top-level element. - Incorrect Iteration Over Business: You tried to loop over a string/empty value instead of the actual Business list.
- Missing Handling for Empty Business List: You didn't set
BusinesstoNonewhen the resulting list is empty as required. - Unsafe Key Access: Directly accessing keys like
el['_source']['id']caused KeyError when those keys are missing.
Corrected Code
def do_the_thing(lst): resp = [] parents_mapper = { 'Marketing': 'level1', 'Advertising': 'level2' } for el in lst: # Extract source once to simplify repeated access source = el.get("_source", {}) # Initialize result dict with safe key access to avoid KeyError d = { 'id': source.get("id"), 'name': source.get("name"), 'Business': [], 'level1': [], 'level2': [] } # Process Business names business_list = source.get("Business", []) for business in business_list: business_name = business.get('name') if business_name: # Only add if name exists and is truthy d['Business'].append(business_name) # Set Business to None if the list is empty (per requirement) if not d['Business']: d['Business'] = None # Process policies (now correctly accessing from source) policies_list = source.get("policies", []) for policy in policies_list: policy_parent = policy.get('parent') target_level = parents_mapper.get(policy_parent) policy_name = policy.get('name') # Only add if parent is recognized and name exists if target_level and policy_name: d[target_level].append(policy_name) resp.append(d) return resp if __name__ == "__main__": import pprint pp = pprint.PrettyPrinter(4) # First test case searchtest1 = [{'_index': 'newtest', '_type': '_doc', '_id': '100', '_score': 1.0, '_source': {'id': '100', 'name': 'A', 'Business': [{'id': '7', 'name': 'Enterprise'}, {'id': '8', 'name': 'Customer'}], 'policies': [{'id': '332', 'name': 'Second division', 'parent': 'Marketing'}, {'id': '3323', 'name': 'First division', 'parent': 'Marketing'}]}}, {'_index': 'newtest', '_type': '_doc', '_id': '101', '_score': 1.0, '_source': {'id': '101', 'name': 'B', 'Business': [{'id': '7'}, {'id': '8', 'name': 'Customer'}], 'policies': [{'id': '332', 'name': 'Second division', 'parent': 'Marketing'}, {'id': '3323', 'name': 'First division', 'parent': 'Marketing'}]}}] print("First Test Case Output:") pp.pprint(do_the_thing(searchtest1)) # Second test case (with missing Business and partial Business entries) searchtest2 = [{'_index': 'newtest', '_type': '_doc', '_id': '100', '_score': 1.0, '_source': {'id': '100', 'name': 'A', 'policies': [{'id': '332', 'name': 'Second division', 'parent': 'Marketing'}, {'id': '3323', 'name': 'First division', 'parent': 'Marketing'}]}}, {'_index': 'newtest', '_type': '_doc', '_id': '101', '_score': 1.0, '_source': {'id': '101', 'name': 'B', 'Business': [{'id': '9'}, {'id': '10', 'name': 'Customer'}], 'policies': [{'id': '332', 'name': 'Second division', 'parent': 'Marketing'}, {'id': '3323', 'name': 'First division', 'parent': 'Advertising'}]}}] print("\nSecond Test Case Output:") pp.pprint(do_the_thing(searchtest2))
Expected Output for First Test Case
[ { "id": "100", "name": "A", "Business": ["Enterprise", "Customer"], "level1": ['Second division', 'First division'], "level2": [] }, { "id": "101", "name": "B", "Business": ["Customer"], "level1": ['Second division', 'First division'], "level2": [] } ]
(Note: Your original expected output had ["Enterprise", "Customer"] for the second entry, but since that entry's first Business item lacks a name, it should only include "Customer" — this is the correct behavior per your requirements.)
Expected Output for Second Test Case
[ { "id": "100", "name": "A", "Business": None, "level1": ['Second division', 'First division'], "level2": [] }, { "id": "101", "name": "B", "Business": ["Customer"], "level1": ['Second division'], "level2": ['First division'] } ]
Key Fixes Explained
- Safe Key Access: Using
.get()everywhere prevents KeyError when keys are missing, returning a default value (likeNoneor empty list) instead. - Correct Business Handling: We now properly iterate over the
Businesslist, collect valid names, and set the value toNoneif no names are found. - Policy Path Correction: Policies are now accessed from the
_sourcedict where they belong, so level1/level2 are populated correctly. - Robust Data Filtering: Only valid policy names (with recognized parents) and Business names are added to the result.
内容的提问来源于stack exchange,提问作者aysh
相关产品推荐
相关产品推荐

