You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python:删除指定列(salary、age)中值为0的行

问题回顾

我们有一个包含薪资、年龄和性别的DataFrame,其中salary和age列的0属于缺失数据,但gender列的0是有效标识(代表女性)。需要删除所有salary或age为0的行,得到干净的数据集。

原始数据如下:

>>> df
   salary  age  gender
0   10000   23       1
1   15000   34       0
2   23000   21       1
3       0   20       0
4   28500    0       1
5   35000   37       1

目标输出:

>>> df
   salary  age  gender
0   10000   23       1
1   15000   34       0
2   23000   21       1
3   35000   37       1
解决方案

方法一:布尔索引直接过滤

这是最直观的方式,直接筛选出salary和age都不为0的行:

# 筛选符合条件的行
df_cleaned = df[(df['salary'] != 0) & (df['age'] != 0)]
# 可选:重置索引为连续整数
df_cleaned = df_cleaned.reset_index(drop=True)

解释:

  • (df['salary'] != 0) & (df['age'] != 0) 生成一个布尔数组,只有当两列都不为0时才返回True
  • 用这个数组作为掩码筛选DataFrame,就能保留符合要求的行
  • reset_index(drop=True) 是可选操作,用来移除原索引、生成从0开始的新索引,和目标输出格式一致

方法二:将0转为缺失值后删除

这种方法更贴合"把0视为缺失数据"的定义,先把指定列的0替换为NaN,再删除包含缺失值的行:

import numpy as np

# 仅将salary和age列的0替换为缺失值
df_with_nan = df.replace({'salary': {0: np.nan}, 'age': {0: np.nan}})
# 删除包含缺失值的行
df_cleaned = df_with_nan.dropna()
# 重置索引
df_cleaned = df_cleaned.reset_index(drop=True)

解释:

  • replace()方法精准替换指定列的0为NaN,不会影响gender列的0
  • dropna()默认删除任何包含缺失值的行,刚好满足我们的需求
验证结果

运行任意一种方法后,打印df_cleaned就能得到目标输出:

print(df_cleaned)
# 输出:
   salary  age  gender
0   10000   23       1
1   15000   34       0
2   23000   21       1
3   35000   37       1

内容的提问来源于stack exchange,提问作者蔡嚴毅

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 03:58:08