You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python实现:识别含100列名的DataFrame异常值并替换为NaN

解决方案

步骤说明

  1. 识别目标列:筛选出DataFrame中列名包含"100"的列;
  2. 类型转换:将这些列的字符串值转换为数值类型,确保能进行大小比较;
  3. 替换异常值:对数值大于100或小于0的元素,替换为NaN。

完整代码实现

import pandas as pd
import numpy as np

def clean_100_columns(df):
    # 筛选列名包含"100"的列
    target_columns = df.columns[df.columns.str.contains('100')]
    
    for col in target_columns:
        # 将列转换为数值类型,无法转换的自动设为NaN
        df[col] = pd.to_numeric(df[col], errors='coerce')
        # 替换大于100或小于0的值为NaN
        df[col] = df[col].mask((df[col] > 100) | (df[col] < 0))
    
    return df

# 测试示例
data = {'first_100':  ['25', '1568200', '5'],
        'second_column': ['first_value', 'second_value', 'third_value'],
        'third_100':['89', '9', '589'],
        'fourth_column':['first_value', 'second_value', 'third_value'],
       }

df = pd.DataFrame(data)
cleaned_df = clean_100_columns(df)
print(cleaned_df)

输出结果

first_100 second_column  third_100 fourth_column
0       25.0   first_value       89.0   first_value
1        NaN  second_value        9.0  second_value
2        5.0   third_value        NaN   third_value

代码解释

  • 筛选目标列:df.columns.str.contains('100')生成布尔数组,用于过滤出符合命名规则的列;
  • 类型转换:pd.to_numeric处理字符串转数值,errors='coerce'参数保证非数字内容转为NaN,避免报错;
  • 替换异常值:mask方法会将满足(值>100 或 值<0)条件的元素替换为NaN,逻辑直观易读。

内容的提问来源于stack exchange,提问作者yoopiyo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.19 05:45:42