如何重塑Pandas DataFrame?城市指标长表转宽表需求
解决长格式DataFrame转宽格式的问题
这是典型的长表转宽表需求,用Pandas的pivot()方法就能完美解决你的问题,因为你的数据中每个city+indicator的组合都是唯一的,不需要额外的聚合操作。
完整代码示例
import pandas as pd # 先构造你的原始数据集(如果已经从文件加载好可以跳过这步) raw_data = { 'city': ['Tokio', 'Boston', 'London', 'Tokio', 'Boston', 'London', 'Tokio', 'Boston', 'London', 'Tokio', 'Boston', 'London'], 'indicator': ['Solid Waste Recycled', 'Solid Waste Recycled', 'Solid Waste Recycled', 'Own-Source Revenues', 'Own-Source Revenues', 'Own-Source Revenues', 'Green Area', 'Green Area', 'Green Area', 'City Land Area', 'City Land Area', 'City Land Area'], 'Value': [1.162e+01, 3.912e+01, 0.0, 1.42e+00, 0.0, 3.247e+01, 4.3031e+02, 7.16635e+01, 1.99761e+01, 9.91e+01, 4.2e+01, 8.956e+01] } df = pd.DataFrame(raw_data) # 核心:将长表转为宽表 df1 = df.pivot( index='city', # 用city作为行索引 columns='indicator',# 用indicator的不同值作为列名 values='Value' # 填充单元格的数值来源 ) # 可选:去掉列索引的名称(让输出格式和你想要的完全一致) df1.columns.name = None # 查看结果 print(df1)
输出结果
Solid Waste Recycled Own-Source Revenues Green Area City Land Area Tokio 11.62 1.42 430.310 99.1 Boston 39.12 0.00 71.663 42.0 London 0.00 32.47 19.976 89.56
补充说明
如果你的数据中存在同一个城市+指标有多条记录的情况(比如重复数据),那就要用pivot_table()并指定聚合函数,例如:
# 假设存在重复,取平均值 df1 = df.pivot_table( index='city', columns='indicator', values='Value', aggfunc='mean' # 可以换成sum、max等 )
内容的提问来源于stack exchange,提问作者emax
相关产品推荐
相关产品推荐

