如何用Pandas修改DataFrame:将出现≥2次的列值替换为Other
Pandas实现按列替换高频值为"Other"
原始DataFrame
IP Routing Banking 1 1 6 2 1 6 3 1 7 3 3 8 4 5 9 5 9 7
需求说明
对DataFrame的每一列,若某个值出现2次及以上,就将该值替换为"Other"。
期望输出
IP Routing Banking 1 Other Other 2 Other Other Other Other Other Other 3 8 4 5 9 5 9 Other
实现代码
首先构造原始DataFrame:
import pandas as pd df = pd.DataFrame({ 'IP': [1, 2, 3, 3, 4, 5], 'Routing': [1, 1, 1, 3, 5, 9], 'Banking': [6, 6, 7, 8, 9, 7] })
接着定义替换逻辑并应用到每一列:
def replace_high_freq_values(col): # 统计列内各值的出现次数 value_counts = col.value_counts() # 筛选出出现次数≥2的值 target_values = value_counts[value_counts >= 2].index # 替换目标值为"Other" return col.replace(target_values, 'Other') # 对所有列应用替换逻辑 result_df = df.apply(replace_high_freq_values) # 查看结果 print(result_df)
代码说明
value_counts():统计列中每个唯一值的出现次数,返回一个以值为索引、次数为值的Series。- 筛选
value_counts >= 2的索引,得到需要替换的目标值集合。 replace():将列中所有属于目标值集合的元素替换为"Other"。apply():将上述逻辑批量应用到DataFrame的每一列。
内容的提问来源于stack exchange,提问作者Eisen
相关产品推荐
相关产品推荐

