You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Pivot pandas DataFrame时避免引入NaN值

问题:Pivot操作后DataFrame出现NaN值,如何生成无NaN的目标格式?

我有一个DataFrame,执行pivot操作后行中意外出现了NaN值。

原始代码:

import pandas as pd

category = ['animal', 'animal', 'animal', 'fruit', 'fruit', 'fruit', 'veggie', 'veggie', 'veggie']
obj = ['animal_1', 'animal_2', 'animal_3', 'fruit_1', 'fruit_2', 'fruit_3', 'veggie_1', 'veggie_2', 'veggie_3']

df = pd.DataFrame(list(zip(category, obj)), columns=['category', 'object'])

# pivot操作
result = df.pivot(columns='category', values='object')

我需要得到如下代码的输出效果,关键是结果中不能包含任何NaN值:

animals=['animal_1', 'animal_2', 'animal_3']
fruit=['fruit_1', 'fruit_2', 'fruit_3']
veggies=['veggie_1', 'veggie_2', 'veggie_3']

pd.DataFrame(list(zip(animals, fruit, veggies)), columns=['animal', 'fruit', 'veggie'])
解决方案

问题出在原始pivot操作没有指定行索引,每个分类的条目分散在不同的原始行中,导致pivot后各列无法对齐,产生NaN。解决思路是给每个分类的条目添加组内序号,让同位置的条目共享同一个行索引,再执行pivot。

实现代码:

import pandas as pd

category = ['animal', 'animal', 'animal', 'fruit', 'fruit', 'fruit', 'veggie', 'veggie', 'veggie']
obj = ['animal_1', 'animal_2', 'animal_3', 'fruit_1', 'fruit_2', 'fruit_3', 'veggie_1', 'veggie_2', 'veggie_3']

df = pd.DataFrame(list(zip(category, obj)), columns=['category', 'object'])
# 给每个分类的条目添加组内序号(从0开始计数)
df['group_idx'] = df.groupby('category').cumcount()
# 以组内序号为行索引执行pivot,最后移除多余的索引列
result = df.pivot(index='group_idx', columns='category', values='object').reset_index(drop=True)

print(result)

输出结果与目标格式完全一致,无任何NaN值:

animal    fruit    veggie
0  animal_1  fruit_1  veggie_1
1  animal_2  fruit_2  veggie_2
2  animal_3  fruit_3  veggie_3

内容的提问来源于stack exchange,提问作者psychcoder

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.11 15:31:17