You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Pandas修改DataFrame:将出现≥2次的列值替换为Other

Pandas实现按列替换高频值为"Other"

原始DataFrame

IP Routing Banking
1  1        6
2  1        6
3  1        7
3  3        8
4  5        9
5  9        7

需求说明

对DataFrame的每一列,若某个值出现2次及以上,就将该值替换为"Other"。

期望输出

IP       Routing      Banking
1        Other        Other
2        Other        Other
Other    Other        Other
Other    3            8
4        5            9
5        9            Other

实现代码

首先构造原始DataFrame:

import pandas as pd

df = pd.DataFrame({
    'IP': [1, 2, 3, 3, 4, 5],
    'Routing': [1, 1, 1, 3, 5, 9],
    'Banking': [6, 6, 7, 8, 9, 7]
})

接着定义替换逻辑并应用到每一列:

def replace_high_freq_values(col):
    # 统计列内各值的出现次数
    value_counts = col.value_counts()
    # 筛选出出现次数≥2的值
    target_values = value_counts[value_counts >= 2].index
    # 替换目标值为"Other"
    return col.replace(target_values, 'Other')

# 对所有列应用替换逻辑
result_df = df.apply(replace_high_freq_values)

# 查看结果
print(result_df)

代码说明

  1. value_counts():统计列中每个唯一值的出现次数,返回一个以值为索引、次数为值的Series。
  2. 筛选value_counts >= 2的索引,得到需要替换的目标值集合。
  3. replace():将列中所有属于目标值集合的元素替换为"Other"。
  4. apply():将上述逻辑批量应用到DataFrame的每一列。

内容的提问来源于stack exchange,提问作者Eisen

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.20 18:39:41