You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

嵌套字典元素提取问题:Python代码修复与键缺失处理

Fixing KeyError and Empty List Issues in Your Python Data Processing Code

Let's break down the issues in your existing code and fix them step by step:

Key Issues Identified

  1. Incorrect Key Case for Business: You used lowercase business instead of the correct capitalized Business when accessing the list.
  2. Wrong Path for Policies: Policies are nested inside _source, but your code was looking for them directly in the top-level element.
  3. Incorrect Iteration Over Business: You tried to loop over a string/empty value instead of the actual Business list.
  4. Missing Handling for Empty Business List: You didn't set Business to None when the resulting list is empty as required.
  5. Unsafe Key Access: Directly accessing keys like el['_source']['id'] caused KeyError when those keys are missing.

Corrected Code

def do_the_thing(lst):
    resp = []
    parents_mapper = {
        'Marketing': 'level1',
        'Advertising': 'level2'
    }
    for el in lst:
        # Extract source once to simplify repeated access
        source = el.get("_source", {})
        
        # Initialize result dict with safe key access to avoid KeyError
        d = {
            'id': source.get("id"),
            'name': source.get("name"),
            'Business': [],
            'level1': [],
            'level2': []
        }
        
        # Process Business names
        business_list = source.get("Business", [])
        for business in business_list:
            business_name = business.get('name')
            if business_name:  # Only add if name exists and is truthy
                d['Business'].append(business_name)
        
        # Set Business to None if the list is empty (per requirement)
        if not d['Business']:
            d['Business'] = None
        
        # Process policies (now correctly accessing from source)
        policies_list = source.get("policies", [])
        for policy in policies_list:
            policy_parent = policy.get('parent')
            target_level = parents_mapper.get(policy_parent)
            policy_name = policy.get('name')
            
            # Only add if parent is recognized and name exists
            if target_level and policy_name:
                d[target_level].append(policy_name)
        
        resp.append(d)
    return resp

if __name__ == "__main__":
    import pprint
    pp = pprint.PrettyPrinter(4)
    
    # First test case
    searchtest1 = [{'_index': 'newtest', '_type': '_doc', '_id': '100', '_score': 1.0, '_source': {'id': '100', 'name': 'A', 'Business': [{'id': '7', 'name': 'Enterprise'}, {'id': '8', 'name': 'Customer'}], 'policies': [{'id': '332', 'name': 'Second division', 'parent': 'Marketing'}, {'id': '3323', 'name': 'First division', 'parent': 'Marketing'}]}}, {'_index': 'newtest', '_type': '_doc', '_id': '101', '_score': 1.0, '_source': {'id': '101', 'name': 'B', 'Business': [{'id': '7'}, {'id': '8', 'name': 'Customer'}], 'policies': [{'id': '332', 'name': 'Second division', 'parent': 'Marketing'}, {'id': '3323', 'name': 'First division', 'parent': 'Marketing'}]}}]
    print("First Test Case Output:")
    pp.pprint(do_the_thing(searchtest1))
    
    # Second test case (with missing Business and partial Business entries)
    searchtest2 = [{'_index': 'newtest', '_type': '_doc', '_id': '100', '_score': 1.0, '_source': {'id': '100', 'name': 'A', 'policies': [{'id': '332', 'name': 'Second division', 'parent': 'Marketing'}, {'id': '3323', 'name': 'First division', 'parent': 'Marketing'}]}}, {'_index': 'newtest', '_type': '_doc', '_id': '101', '_score': 1.0, '_source': {'id': '101', 'name': 'B', 'Business': [{'id': '9'}, {'id': '10', 'name': 'Customer'}], 'policies': [{'id': '332', 'name': 'Second division', 'parent': 'Marketing'}, {'id': '3323', 'name': 'First division', 'parent': 'Advertising'}]}}]
    print("\nSecond Test Case Output:")
    pp.pprint(do_the_thing(searchtest2))

Expected Output for First Test Case

[ { "id": "100", "name": "A", "Business": ["Enterprise", "Customer"], "level1": ['Second division', 'First division'], "level2": [] }, { "id": "101", "name": "B", "Business": ["Customer"], "level1": ['Second division', 'First division'], "level2": [] } ]

(Note: Your original expected output had ["Enterprise", "Customer"] for the second entry, but since that entry's first Business item lacks a name, it should only include "Customer" — this is the correct behavior per your requirements.)

Expected Output for Second Test Case

[ { "id": "100", "name": "A", "Business": None, "level1": ['Second division', 'First division'], "level2": [] }, { "id": "101", "name": "B", "Business": ["Customer"], "level1": ['Second division'], "level2": ['First division'] } ]

Key Fixes Explained

  • Safe Key Access: Using .get() everywhere prevents KeyError when keys are missing, returning a default value (like None or empty list) instead.
  • Correct Business Handling: We now properly iterate over the Business list, collect valid names, and set the value to None if no names are found.
  • Policy Path Correction: Policies are now accessed from the _source dict where they belong, so level1/level2 are populated correctly.
  • Robust Data Filtering: Only valid policy names (with recognized parents) and Business names are added to the result.

内容的提问来源于stack exchange,提问作者aysh

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.08 18:32:26