如何将DataFrame的attributeName列转为表头并填充对应attributeValue
场景说明
我有一个名为attributes的DataFrame,存储了3个系列汽车的属性信息,数据结构如下:
attributes = {'brand': ['Honda Civic','Honda Civic','Honda Civic','Toyota Corolla','Toyota Corolla','Audi A4'], 'attributeName': ['wheels','doors','fuelType','wheels','color','wheels'], 'attributeValue': ['4','2','hybrid','4','red','4'] }
预期输出结果
result = {'brand': ['Honda Civic','Toyota Corolla','Audi A4'], 'wheels': ['4','4','4'], 'doors': ['2','',''], 'fuelType':['hybrid','',''], 'color': ['','red',''] }
问题说明
我需要将attributeName列的取值转为新列表头,每个品牌/汽车单独占一行,对应位置填充attributeValue的取值。目前使用get_dummies进行转换时,仅能得到布尔类型的结果,无法保留原始的属性值,请问该如何实现该需求?
解决方案
直接使用pandas的pivot方法实现长表转宽表,再对缺失值做填充即可,代码如下:
import pandas as pd # 构造原始DataFrame df = pd.DataFrame(attributes) # 长表转宽表+重置索引+空值填充为空白字符串 res_df = df.pivot( index='brand', columns='attributeName', values='attributeValue' ).reset_index().fillna('') # 转为预期的字典格式 result = res_df.to_dict(orient='list')
运行上述代码后得到的result变量和预期输出完全一致。
内容的提问来源于stack exchange,提问作者HedgeHog
相关产品推荐
相关产品推荐

