You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Pandas保留DataFrame重复值首项并将其余设为NaN?

保留DataFrame指定列重复值首次出现,其余设为NaN

示例输入数据

Adescription
Firsthello
Secondhello
Thirdhello
Fourthwhy
Fifthwhy

解决方案

首先构造示例DataFrame(若已有数据可跳过此步骤):

import pandas as pd

df = pd.DataFrame({
    'A': ['First', 'Second', 'Third', 'Fourth', 'Fifth'],
    'description': ['hello', 'hello', 'hello', 'why', 'why']
})

用Pandas内置方法快速实现需求,以下两种方法效果一致:

方法1:使用where()过滤赋值

df['description'] = df['description'].where(~df['description'].duplicated())

方法2:使用loc直接定位重复项赋值

df.loc[df['description'].duplicated(), 'description'] = pd.NA

输出结果

Adescription
Firsthello
SecondNaN
ThirdNaN
Fourthwhy
FifthNaN

说明

  • duplicated()默认标记除首次出现外的所有重复项为True;
  • where()会保留布尔条件为True的原数据,将False位置替换为NaN;
  • 两种方法逻辑等价,可根据个人习惯选择。

内容的提问来源于stack exchange,提问作者glitch_123

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.20 08:21:57