使用fillna()后Pandas DataFrame仍存在NaN值的问题排查
问题:Forward Fill后仍存在NaN值排查
原始数据与代码
创建DataFrame
import numpy as np import pandas as pd df = pd.DataFrame({'open':[0, 1, 2, 3, 4], 'close':[1, 2, 2, 3, 2]}) print(df)
输出:
open close 0 0 1 1 1 2 2 2 2 3 3 3 4 4 2
自定义分类函数
def classify_color(df: pd.DataFrame): conditions = [df.close < df.open, df.close > df.open, df.close == df.open] choices = ["Red", "Green", "Grey"] res = np.select(condlist=conditions, choicelist=choices) # 分类为Red/Green/Grey res = np.where(res == "Grey", np.nan, res) # 将Grey替换为NaN res = pd.DataFrame({"C": res}).fillna(method="ffill") # 向前填充NaN res = res.C.values df['candle_color'] = res
调用后异常结果
classify_color(df=df) print(df.candle_color.value_counts())
输出:
Green 2 nan 2 Red 1 Name: candle_color, dtype: int64
问题原因
核心问题是**fillna(method="ffill")在新版Pandas(2.1.0+)中已被弃用**,这个写法不会执行填充操作,导致原本的NaN保留了下来。
从数据逻辑看,原始数据中索引2、3的close等于open,会被转为NaN;索引0是Green、索引1是Green、索引4是Red,正常填充后这两个NaN应该被填充为Green,但由于填充方法失效,NaN仍存在。
修复方案
方案1:使用Pandas推荐的ffill()方法
直接替换fillna(method="ffill")为ffill(),这是当前版本的标准写法:
def classify_color(df: pd.DataFrame): conditions = [df.close < df.open, df.close > df.open, df.close == df.open] choices = ["Red", "Green", "Grey"] res = np.select(condlist=conditions, choicelist=choices) res = np.where(res == "Grey", np.nan, res) # 替换为ffill()方法 res = pd.DataFrame({"C": res}).ffill().C.values df['candle_color'] = res
方案2:直接在numpy数组上实现向前填充(更高效)
不需要转成DataFrame,用numpy实现ffill逻辑,避免版本兼容问题:
def classify_color(df: pd.DataFrame): conditions = [df.close < df.open, df.close > df.open, df.close == df.open] choices = ["Red", "Green", "Grey"] res = np.select(condlist=conditions, choicelist=choices) res = np.where(res == "Grey", np.nan, res) # numpy实现向前填充 mask = pd.notna(res) idx = np.where(mask, np.arange(len(res)), 0) np.maximum.accumulate(idx, out=idx) res = res[idx] df['candle_color'] = res
修复后结果
调用函数后再查看value_counts:
Green 4 Red 1 Name: candle_color, dtype: int64
内容的提问来源于stack exchange,提问作者jamesB
相关产品推荐
相关产品推荐

