You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用pd.pivot_table或groupby获取子组唯一值列表并保留为DataFrame

解决方案

以下两种方法均无需循环,输出直接为DataFrame格式,可实现需求:

方法1:修改pivot_table聚合函数

你原有代码用sum做聚合会直接拼接分组内的字符串,把聚合函数替换为「取唯一值后转列表」即可:

import pandas as pd

df = pd.DataFrame({
    'number': ['101','101','101','201','101','101','201','201'], 
    'letter': ['a','b','a','b','b','a','b','a'], 
    'fruit': ['apple','melon','peach','grape','orange','pear','apple','peach'] 
})

result = pd.pivot_table(
    df, 
    index='number', 
    columns='letter', 
    values='fruit', 
    aggfunc=lambda x: list(x.unique())
)

方法2:groupby+unstack实现(写法更简洁)

result = df.groupby(['number', 'letter'])['fruit'].unique().apply(list).unstack()

如果需要对分组内的列表元素排序,把list(x.unique())替换为sorted(x.unique())即可。
运行后输出结果如下:

numberab
101['apple', 'peach', 'pear']['melon', 'orange']
201['peach']['grape', 'apple']

注:你给出的期望输出中101对应a列少了peach属于笔误,实际运行会返回所有唯一值,如有特殊过滤需求可在聚合函数中额外加判断条件。

内容的提问来源于stack exchange,提问作者j__carlson

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.02 11:09:04