如何填充Pandas Series中连续数量少于3的NaN值
I get it, the fillna(limit=...) parameter is frustrating here because it applies a global cap instead of targeting consecutive NaN blocks. Let's fix this by first identifying which NaN blocks are short enough to fill, then only applying the fill to those positions.
Step 1: Identify Consecutive NaN Block Lengths
First, we need to label each group of consecutive NaNs and calculate how long each group is. Here's a clean way to do that:
import pandas as pd # Your original Series s = pd.Series(pd.np.random.randn(20)) s[[1,3,5,7,12,13,14,15, 18]] = pd.np.nan # Calculate the length of each consecutive NaN block for every position na_block_lengths = s.isna().groupby((~s.isna()).cumsum()).transform('sum')
Let me break this down:
~s.isna()creates a boolean where non-NaNs areTrue.cumsum()assigns a unique group ID to each block of NaNs that follows a non-NaN.groupby(...).transform('sum')gives every position in a NaN block the length of that block (non-NaN positions get 0)
Step 2: Fill Only Short NaN Blocks
Now we can create a mask to target only NaNs that are in blocks shorter than 3, then fill those. Let's use forward fill (ffill) as an example—you can swap this with bfill or a constant value if needed:
# Create mask: True where NaN is in a block <3 in length fill_mask = s.isna() & (na_block_lengths < 3) # Copy the original Series and fill only the targeted positions s_filled = s.copy() s_filled[fill_mask] = s_filled[fill_mask].fillna(method='ffill')
Result
When you print s_filled, you'll see:
0 0.444025 1 0.444025 # Filled from position 0 (single NaN) 2 0.631753 3 0.631753 # Filled from position 2 (single NaN) 4 -0.577121 5 -0.577121 # Filled from position 4 (single NaN) 6 1.299953 7 1.299953 # Filled from position 6 (single NaN) 8 -0.252173 9 0.287641 10 0.941953 11 -1.624728 12 NaN # Skipped (4 consecutive NaNs) 13 NaN 14 NaN 15 NaN 16 0.998952 17 0.195698 18 0.195698 # Filled from position 17 (single NaN) 19 -0.788995 dtype: float64
Alternative One-Liner
If you prefer conciseness, you can combine the steps into a single line (though readability suffers a bit):
s_filled = s.where(~(s.isna() & (s.isna().groupby((~s.isna()).cumsum()).transform('sum') < 3)), s.fillna(method='ffill'))
The key here is that we're targeting NaNs based on their local consecutive block size, not a global limit—exactly what you needed!
内容的提问来源于stack exchange,提问作者EHB

