Python实现:识别含100列名的DataFrame异常值并替换为NaN
解决方案
步骤说明
- 识别目标列:筛选出DataFrame中列名包含"100"的列;
- 类型转换:将这些列的字符串值转换为数值类型,确保能进行大小比较;
- 替换异常值:对数值大于100或小于0的元素,替换为
NaN。
完整代码实现
import pandas as pd import numpy as np def clean_100_columns(df): # 筛选列名包含"100"的列 target_columns = df.columns[df.columns.str.contains('100')] for col in target_columns: # 将列转换为数值类型,无法转换的自动设为NaN df[col] = pd.to_numeric(df[col], errors='coerce') # 替换大于100或小于0的值为NaN df[col] = df[col].mask((df[col] > 100) | (df[col] < 0)) return df # 测试示例 data = {'first_100': ['25', '1568200', '5'], 'second_column': ['first_value', 'second_value', 'third_value'], 'third_100':['89', '9', '589'], 'fourth_column':['first_value', 'second_value', 'third_value'], } df = pd.DataFrame(data) cleaned_df = clean_100_columns(df) print(cleaned_df)
输出结果
first_100 second_column third_100 fourth_column 0 25.0 first_value 89.0 first_value 1 NaN second_value 9.0 second_value 2 5.0 third_value NaN third_value
代码解释
- 筛选目标列:
df.columns.str.contains('100')生成布尔数组,用于过滤出符合命名规则的列; - 类型转换:
pd.to_numeric处理字符串转数值,errors='coerce'参数保证非数字内容转为NaN,避免报错; - 替换异常值:
mask方法会将满足(值>100 或 值<0)条件的元素替换为NaN,逻辑直观易读。
内容的提问来源于stack exchange,提问作者yoopiyo
相关产品推荐
相关产品推荐

