You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何按唯一ID透视pandas DataFrame并正确保留二值列取值

问题原因

你之前的pivot_table写法参数设置错误:将place设为行索引、ID设为列索引,本质是做两个维度的交叉透视,和「每个唯一ID对应一行、按字段类型分别聚合」的需求完全不匹配,因此输出结果和预期差异极大。

本次聚合的核心规则分两类:

  • 固定属性字段(place、sex):同一ID下取值完全一致,聚合时保留该唯一值即可
  • 二值标识字段(depression、stressed、sleep、ate):同一ID分组内只要存在取值1,聚合结果就为1,等价于取分组内的最大值
实现方案

最直接的实现方式是用groupby配合自定义聚合规则,代码如下:

import pandas as pd

# 原始示例数据
example = {
"ID": [1, 1, 2, 2, 2, 3],
"place":["Maryland","Maryland", "Washington", "Washington", "Washington", "Los Angeles"],
"sex":["male","male","female", "female", "female", "other"],
"depression": [0, 0, 0, 0, 0, 1],
"stressed":  [1 ,0, 0, 0, 0, 0],
"sleep": [1, 1, 1, 0, 1, 1],
"ate":[0,1, 0, 1, 0, 1],
}
example = pd.DataFrame(example)

# 为不同字段指定对应的聚合函数
agg_rule = {
    "place": "first",
    "sex": "first",
    "depression": "max",
    "stressed": "max",
    "sleep": "max",
    "ate": "max"
}

# 按ID分组聚合,as_index=False保证ID作为普通列输出而非行索引
result = example.groupby("ID", as_index=False).agg(agg_rule)
print(result)

如果要使用pivot_table实现,只需要正确指定行索引和聚合规则即可,参考代码:

result = example.pivot_table(
    index="ID",
    aggfunc=agg_rule,
    as_index=False
)
输出结果

上述代码运行后输出和你给出的目标格式完全一致:

ID        place     sex  depression  stressed  sleep  ate
0   1     Maryland    male           0         1      1    1
1   2   Washington  female           0         0      1    1
2   3  Los Angeles   other           1         0      1    1

注:如果后续固定属性列存在同一ID下多个不同取值的异常情况,first只会取分组内第一个值,可以替换为lambda x: x.unique()[0]做校验,或者根据业务需求调整取值逻辑。

内容的提问来源于stack exchange,提问作者Shu

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.29 02:33:08