如何从Pandas字典列提取元素生成新列?遇KeyError报错
Pandas字典列批量提取元素报错的解决方法
问题场景
你有一个包含ID和data(字典类型)两列的DataFrame:
ID data 0 6602629924 {'@status': 'found', '@_fa': 'true', 'coredata... 1 55599317400 {'@status': 'found', '@_fa': 'true', 'coredata... 2 25652391600 {'@status': 'found', '@_fa': 'true', 'coredata... 3 11939875400 {'@status': 'found', '@_fa': 'true', 'coredata... 4 56140547500 {'@status': 'found', '@_fa': 'true', 'coredata...
单行提取affiliation的代码可以正常返回结果:
data[1]["author-profile"]["affiliation-current"]["affiliation"]["ip-doc"]["afdispname"] # 返回 'De La Salle University'
但批量处理整列时执行以下代码:
new_df["affiliation"] = new_df['data']["author-profile"]["affiliation-current"]["affiliation"]["ip-doc"]["afdispname"] new_df
出现KeyError: 'author-profile'报错。
错误原因
new_df['data']是一个Pandas Series对象,直接对它使用["author-profile"]时,Pandas会把这个字符串当作Series的行索引去查找,而不是对Series中的每个字典元素执行索引操作。因为Series的索引里没有'author-profile'这个值,所以抛出KeyError。
而单行代码里的data[1]是直接取出了单个字典对象,所以可以用字典索引的方式逐层取值。
解决方法
方法1:使用apply+自定义函数(带错误处理)
通过apply对Series中的每个字典元素单独处理,同时加入异常捕获避免个别字典缺失键导致报错:
def extract_affiliation(dict_data): try: return dict_data["author-profile"]["affiliation-current"]["affiliation"]["ip-doc"]["afdispname"] except KeyError: return None # 缺失键时返回None,也可以换成NaN new_df["affiliation"] = new_df['data'].apply(extract_affiliation)
方法2:使用链式str.get(简洁自动处理缺失)
Pandas的Series提供了str.get方法,可以直接对字典类型的元素提取键值,链式调用逐层取值,缺失键时自动返回NaN:
new_df["affiliation"] = new_df['data'].str.get("author-profile")\ .str.get("affiliation-current")\ .str.get("affiliation")\ .str.get("ip-doc")\ .str.get("afdispname")
方法3:使用lambda+字典get方法(极简写法)
用lambda函数结合字典的get方法,逐层取值,每一层都设置默认空字典避免KeyError:
new_df["affiliation"] = new_df['data'].apply( lambda x: x.get("author-profile", {}) .get("affiliation-current", {}) .get("affiliation", {}) .get("ip-doc", {}) .get("afdispname") )
内容的提问来源于stack exchange,提问作者Ken.WS
相关产品推荐
相关产品推荐

