Python Pandas:含100列名的DataFrame异常值替换为NaN并关联列值
Python函数实现DataFrame异常值处理与提取
首先注意:原始数据中third_100和fourth_100列的数值是以字符串形式存储的,必须先转换为数值类型才能进行大小判断,这是实现功能的前提。
以下是完整的实现代码,包含所有要求的功能:
import pandas as pd import numpy as np def process_df_with_100_cols(df): # 1. 筛选列名包含"100"的目标列 target_cols = [col for col in df.columns if '100' in col] # 将目标列转换为数值类型,无法转换的内容自动转为NaN df[target_cols] = df[target_cols].apply(pd.to_numeric, errors='coerce') # 2. 将大于100或小于0的数值替换为NaN df[target_cols] = df[target_cols].mask((df[target_cols] > 100) | (df[target_cols] < 0)) # 3. 获取所有存在异常值的行对应的first_column值 abnormal_row_mask = df[target_cols].isna().any(axis=1) abnormal_first_col_values = df.loc[abnormal_row_mask, 'first_column'].tolist() return df, abnormal_first_col_values # 初始化原始DataFrame data = { 'first_column': ['product_name', 'product_name2', 'product_name3'], 'second_column': ['first_value', 'second_value', 'third_value'], 'third_100': ['89', '9', '589'], 'fourth_100': ['25', '1568200', '5'] } df = pd.DataFrame(data) # 执行处理 processed_df, abnormal_values = process_df_with_100_cols(df) # 输出结果 print("处理后的DataFrame:") print(processed_df) print("\n异常值对应的first_column值:") print(abnormal_values)
代码说明
- 筛选目标列:用列表推导式遍历所有列名,直接筛选包含"100"的列,简单高效。
- 转换数值类型:使用
pd.to_numeric将字符串列转为数值列,errors='coerce'参数确保非数值内容转为NaN,避免类型错误。 - 替换异常值:使用
mask方法,当数值满足>100或<0时替换为NaN,逻辑清晰。 - 提取异常行对应的first_column:通过
isna().any(axis=1)标记存在异常值的行,再用loc提取对应列的值并转为列表。
运行结果
处理后的DataFrame:
first_column second_column third_100 fourth_100 0 product_name first_value 89.0 25.0 1 product_name2 second_value 9.0 NaN 2 product_name3 third_value NaN 5.0
异常值对应的first_column值:
['product_name2', 'product_name3']
内容的提问来源于stack exchange,提问作者yoopiyo
相关产品推荐
相关产品推荐

