如何在Pandas DataFrame中跳过NaN并左移行元素?支持大数据集
First, let's clarify the desired output from your sample DataFrame. Given this input:
import pandas as pd import numpy as np df = pd.DataFrame({ 'A': ['a', 'b', 'c', 'd'], 'B': [np.nan, np.nan, np.nan, 'a'], 'C': ['c', 'b', np.nan, 'b'], 'D': [np.nan, 'a', 'd', 'c'] })
You want to shift all non-NaN values to the left (preserving their original order) and fill remaining positions with NaNs, resulting in:
A B C D 0 a c NaN NaN 1 b b a NaN 2 c d NaN NaN 3 d a b c
Scalable Vectorized Solution (Best for Large Datasets)
Row-wise apply methods work for small data but are way too slow for 100k+ rows. Instead, use vectorized numpy operations—these are optimized for speed and memory efficiency, making them perfect for big datasets:
# Convert DataFrame to numpy array and create non-NaN mask arr = df.values mask = df.notna().values # Get indices that sort non-NaNs to the left (stable sort preserves original order) sorted_indices = np.argsort(~mask, axis=1, kind='mergesort') # Reorder the array using these indices result_arr = arr[np.arange(len(arr))[:, None], sorted_indices] # Convert back to DataFrame with original columns df_result = pd.DataFrame(result_arr, columns=df.columns)
Why This Works:
~maskcreates a boolean array whereTruemarks NaNs. When sorted, theseTruevalues get pushed to the right.kind='mergesort'ensures a stable sort, so non-NaN values keep their original relative order (no shuffling of valid data).- Vectorized operations avoid looping over rows, making this approach orders of magnitude faster than row-wise
applyfor large datasets.
Naive (Slow) Alternative for Reference
If you're wondering why your initial apply attempt might have struggled, here's a common naive approach that works but isn't scalable:
def shift_left(row): non_nan = row.dropna().values return pd.Series(np.pad(non_nan, (0, len(row)-len(non_nan)), mode='constant', constant_values=np.nan)) # Works but is slow for 100k rows df_result_slow = df.apply(shift_left, axis=1)
This method iterates over every row individually, which is inefficient for big data. Stick to the vectorized numpy approach for performance.
Performance Note
For a 100k-row dataset with 4 columns, the vectorized method finishes in milliseconds, while the apply method takes several seconds (or longer depending on your hardware). This makes the numpy solution ideal for production-scale data.
内容的提问来源于stack exchange,提问作者nOObda

