如何在Pandas中转换归一化后的DataFrame列结构?
将宽格式属性列转换为长格式DataFrame
你需要把这种带编号的属性列(如attributes_0_color、attributes_1_pet)转换为每个属性对应一行的长格式DataFrame,以下是两种可行的实现方法:
方法一:使用melt + pivot
这种方法步骤直观,适合理解数据重塑的过程:
import pandas as pd # 构造你的原始DataFrame df = pd.DataFrame({ 'id': [1, 2], 'name': ['John', 'Bill'], 'attributes_0_color': ['red', 'blue'], 'attributes_0_pet': ['dog', 'cat'], 'attributes_1_color': ['purple', 'orange'], 'attributes_1_pet': ['parrot', 'hamster'] }) # 1. 将属性列转为长格式,保留id和name作为标识 melted = df.melt(id_vars=['id', 'name'], var_name='attr_col', value_name='value') # 2. 拆分属性列名,提取编号和字段类型(color/pet) melted[['_', 'attr_idx', 'field']] = melted['attr_col'].str.split('_', n=2, expand=True) melted = melted.drop(columns=['_', 'attr_col']) # 3. 重塑为目标格式 result = melted.pivot( index=['id', 'name', 'attr_idx'], columns='field', values='value' ).reset_index().drop(columns='attr_idx').sort_values('id').reset_index(drop=True) print(result)
输出结果:
id name color pet 0 1 John red dog 1 1 John purple parrot 2 2 Bill blue cat 3 2 Bill orange hamster
方法二:使用pd.wide_to_long
这个方法专门针对带索引编号的宽格式列,代码更简洁,处理大数据效率更高:
import pandas as pd # 构造原始DataFrame df = pd.DataFrame({ 'id': [1, 2], 'name': ['John', 'Bill'], 'attributes_0_color': ['red', 'blue'], 'attributes_0_pet': ['dog', 'cat'], 'attributes_1_color': ['purple', 'orange'], 'attributes_1_pet': ['parrot', 'hamster'] }) # 重命名列,适配wide_to_long的格式要求:{字段名}_{编号} df_renamed = df.rename(columns={ 'attributes_0_color': 'color_0', 'attributes_0_pet': 'pet_0', 'attributes_1_color': 'color_1', 'attributes_1_pet': 'pet_1' }) # 执行转换 result = pd.wide_to_long( df_renamed, stubnames=['color', 'pet'], i=['id', 'name'], j='attr_idx', sep='_' ).reset_index().drop(columns='attr_idx').sort_values('id').reset_index(drop=True) print(result)
输出结果和方法一完全一致。
内容的提问来源于stack exchange,提问作者kunshabba
相关产品推荐
相关产品推荐

