如何在Pandas中将含attributes后缀列的字典合并至attributes列
问题描述
现有如下DataFrame:
import pandas as pd df = pd.DataFrame({'id' : [1,2,3], 'attributes' : [{'dd' : True, 'budget' : '35k'}, {'dd' : True, 'budget' : '25k'}, {'dd' : True, 'budget' : '40k'}], 'prod.attributes' : [{'img' : 'img1.url', 'name' : 'millennials'}, {'img' : 'img2.url', 'name' : 'single'}, {'img' : 'img3.url', 'name' : 'married'}]})
输出的df如下:
id attributes prod.attributes 0 1 {'dd': True, 'budget': '35k'} {'img': 'img1.url', 'name': 'millennials'} 1 2 {'dd': True, 'budget': '25k'} {'img': 'img2.url', 'name': 'single'} 2 3 {'dd': True, 'budget': '40k'} {'img': 'img3.url', 'name': 'married'}
需要将所有以attributes为后缀的列的字典数据,以对应前缀为键合并到attributes列中,期望结果如下:
op = pd.DataFrame({'id' : [1,2,3], 'attributes' : [{'dd' : True, 'budget' : '35k', 'prod' : {'img' : 'img1.url', 'name' : 'millennials'}}, {'dd' : True, 'budget' : '25k', 'prod' : {'img' : 'img2.url', 'name' : 'single'}}, {'dd' : True, 'budget' : '40k', 'prod' : {'img' : 'img3.url', 'name' : 'married'}}]})
输出的op如下:
id attributes 0 1 {'dd': True, 'budget': '35k', 'prod': {'img': 'img1.url', 'name': 'millennials'}} 1 2 {'dd': True, 'budget': '25k', 'prod': {'img': 'img2.url', 'name': 'single'}} 2 3 {'dd': True, 'budget': '40k', 'prod': {'img': 'img3.url', 'name': 'married'}}
尝试了以下代码,但返回结果全为None:
df['attributes'].apply(lambda x : x.update({'audience' : df['prod.attributes']}))
错误原因分析
dict.update()方法是原地修改字典,执行后返回None,所以用apply会得到全None的结果。- 代码里直接引用
df['prod.attributes']是整个Series,不是当前行对应的字典值,会把整个列的内容塞到每个字典里,逻辑错误。
解决方案
方法1:处理单列情况
如果只需要处理prod.attributes这一列,可以用apply时传入行数据,逐行合并:
df['attributes'] = df.apply(lambda row: {**row['attributes'], 'prod': row['prod.attributes']}, axis=1) # 之后可以删除原列 df = df.drop('prod.attributes', axis=1)
方法2:通用处理所有以attributes为后缀的列
如果有多个类似xxx.attributes的列,用以下代码自动识别并合并:
# 获取所有以attributes结尾的列,排除原attributes列 attr_cols = [col for col in df.columns if col.endswith('.attributes') and col != 'attributes'] # 逐行合并字典 df['attributes'] = df.apply(lambda row: { **row['attributes'], **{col.split('.')[0]: row[col] for col in attr_cols} }, axis=1) # 删除原有的xxx.attributes列 df = df.drop(attr_cols, axis=1)
执行后就能得到期望的结果,这种方法可以自动适配任意数量的xxx.attributes列,无需手动逐个指定。
内容的提问来源于stack exchange,提问作者Karthik S
相关产品推荐
相关产品推荐

