使用Python reduce函数填充DataFrame的NaN值时出错的原因咨询
Hey there! Let's break down exactly why your reduce + lambda code threw that error, and how you could fix it if you ever wanted to use reduce for this task (though your final solution is way better, by the way!).
The core issue is how Python's reduce function works: it expects the first argument (your lambda) to accept two positional parameters, not one.
Here's what was happening with your code:
reduce(lambda x : df.fillna({x : 0}, inplace=True), ['Col_A', 'Col_B', 'Col_C', 'Col_D'])
reducestarts by taking the first element in your list ('Col_A') as the initial "accumulator" value.- Next, it tries to pass two values to your lambda: the accumulator (
'Col_A') and the next element in the list ('Col_B'). - But your lambda only accepts one parameter (
x), hence the error:<lambda>() takes 1 positional argument but 2 were given.
On top of that, using inplace=True makes this even worse: fillna(inplace=True) modifies your DataFrame in place and returns None. So even if you fixed the lambda parameters, after the first iteration your accumulator would become None, and subsequent calls would fail entirely.
reduce Approach (If You Really Wanted To) If you were set on using reduce, you'd need to adjust the lambda to accept two parameters, and avoid inplace=True so you can pass the updated DataFrame through each iteration. Here's what that would look like:
from functools import reduce # Define your target columns cols = ['Col_A', 'Col_B', 'Col_C', 'Col_D'] # Use reduce with an initial value (your original df) df = reduce(lambda acc_df, col: acc_df.fillna({col: 0}), cols, df)
- The third argument to
reduce(df) sets the initial accumulator value to your original DataFrame. - Each iteration passes the updated DataFrame (
acc_df) and the current column name (col) to the lambda. fillnareturns a new DataFrame with the NaNs replaced for that column, which becomes the accumulator for the next step.
Your fix:
df[['Col_A', 'Col_B', 'Col_C', 'Col_D']] = df[['Col_A', 'Col_B', 'Col_C', 'Col_D']].fillna(0)
is far more idiomatic Pandas. It leverages Pandas' vectorized operations to replace NaNs across all target columns in one go, which is faster and cleaner than iterating through columns with reduce. This is exactly the kind of approach you should reach for first in Pandas.
内容的提问来源于stack exchange,提问作者JBVasc

