You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何仅对DataFrame重复行排名且排除NaN值?

问题描述

现有如下数据表格:

Col1
0   1.0
1   1.0
2   1.0
3   2.0
4   3.0
5   4.0
6   NaN

需求是仅对重复值进行排名(忽略NaN值),但当前代码输出中唯一值也被赋予了排名,不符合要求:

Col1   Rn
0   1.0  1.0
1   1.0  2.0
2   1.0  3.0
3   2.0  1.0
4   3.0  1.0
5   4.0  1.0
6   NaN  NaN

期望的正确输出为:

Col1   Rn
0   1.0  1.0
1   1.0  2.0
2   1.0  3.0
3   2.0  NaN
4   3.0  NaN
5   4.0  NaN
6   NaN  NaN

用户的示例代码如下:

import numpy as np
import pandas as pd

df = pd.DataFrame([[1],
                   [1],
                   [1],
                   [2],
                   [3],
                   [4],
                   [np.NaN]], columns=['Col1'])
print(df)


# Adding row_number for each pair:
df['Rn'] = df[df['Col1'].notnull()].groupby('Col1')['Col1'].rank(method="first", ascending=True)
print(df)

# I managed to select only necessary rows for mask, but how can I apply it along with groupby?:
m = df.dropna().loc[df['Col1'].duplicated(keep=False)]
print(m)
解决方案

可以通过精准标记重复组、仅对目标行执行排名的方式实现需求,修改后的代码如下:

import numpy as np
import pandas as pd

df = pd.DataFrame([[1],
                   [1],
                   [1],
                   [2],
                   [3],
                   [4],
                   [np.NaN]], columns=['Col1'])

# 统计每个值的出现次数,筛选出重复值(出现次数≥2)
value_counts = df['Col1'].value_counts()
duplicate_values = value_counts[value_counts >= 2].index

# 创建掩码:标记需要排名的行(非NaN且属于重复值)
mask = df['Col1'].notnull() & df['Col1'].isin(duplicate_values)

# 仅对掩码覆盖的行执行组内排名,其余行设为NaN
df['Rn'] = np.where(
    mask,
    df[mask].groupby('Col1')['Col1'].rank(method="first", ascending=True),
    np.nan
)

print(df)

代码说明

  • value_counts():统计Col1中各值的出现频次,快速定位重复值
  • duplicate_values:提取出现次数≥2的值,明确需要排名的目标组
  • mask:精准筛选出需要处理的行,排除NaN和唯一值
  • np.where():根据掩码判断,对符合条件的行执行组内排名,不符合的直接设为NaN

运行代码后即可得到期望的输出结果。

内容的提问来源于stack exchange,提问作者muted_buddy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.19 00:27:22