如何避免字典为空或无匹配键时Pandas脚本报错?
Hey there, let's sort out this error you're hitting. The issue almost certainly comes from trying to access a dictionary key that doesn't exist (or working with an empty dictionary) using direct bracket notation like my_dict[row['Key']]—that'll throw a KeyError every time. Let's walk through safe, efficient solutions tailored to your DataFrame.
First, let's start with a corrected version of your DataFrame (I fixed the admit column to match the length of other columns, which would cause an error otherwise):
import pandas as pd df = pd.DataFrame({ 'Key': ['10003', '10003', '10003', '10003', '10003','10003','10034'], 'Num1': [12,13,30,12,13,13,16], 'Num2': [121,122,122,124,125,126,127], 'admit': [2015019, 20150124, 20150206,20150211,20150215,20150220,20150225] })
Solution 1: Safe apply() Handling
If you want to stick with apply(), the key is to check if the key exists in the dictionary first (this automatically handles empty dictionaries too, since key in empty_dict returns False).
Define a helper function that safely checks for the key and returns your marker:
# Example target dictionary (use {} to test empty case) target_dict = {'10003': 'mark_me'} def mark_row(row): # Return '1' if Key is in the dictionary, else '0' return '1' if row['Key'] in target_dict else '0' # Add the new marked column to your DataFrame df['marked_column'] = df.apply(mark_row, axis=1)
This will never throw an error: if the dictionary is empty, all rows get '0'; if a row's Key isn't in the dictionary, it also gets '0'.
Solution 2: Faster Vectorized Operation (Recommended)
apply() is slow for large DataFrames because it processes rows one by one. A better approach is to use pandas' built-in vectorized methods, which are optimized for speed:
# Use isin() to check if Key is in the dictionary's keys, then map to '1'/'0' df['marked_column'] = df['Key'].isin(target_dict.keys()).map({True: '1', False: '0'}) # Alternative: Convert boolean to int then string df['marked_column'] = df['Key'].isin(target_dict.keys()).astype(int).astype(str)
This does the same job but runs orders of magnitude faster on big datasets.
Why Your Original Code Threw Errors
If you were using something like return target_dict[row['Key']] directly, Python will raise a KeyError as soon as it hits a Key that's not in the dictionary (or if the dictionary is empty). By checking row['Key'] in target_dict first, we avoid that entirely.
内容的提问来源于stack exchange,提问作者CandleWax

