You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Pandas DataFrame中跳过NaN并左移行元素?支持大数据集

Efficiently Shift Non-NaN Values Left in Pandas (Scalable to 100k+ Rows)

First, let's clarify the desired output from your sample DataFrame. Given this input:

import pandas as pd
import numpy as np

df = pd.DataFrame({
    'A': ['a', 'b', 'c', 'd'],
    'B': [np.nan, np.nan, np.nan, 'a'],
    'C': ['c', 'b', np.nan, 'b'],
    'D': [np.nan, 'a', 'd', 'c']
})

You want to shift all non-NaN values to the left (preserving their original order) and fill remaining positions with NaNs, resulting in:

A    B    C    D
0  a    c  NaN  NaN
1  b    b    a  NaN
2  c    d  NaN  NaN
3  d    a    b    c

Scalable Vectorized Solution (Best for Large Datasets)

Row-wise apply methods work for small data but are way too slow for 100k+ rows. Instead, use vectorized numpy operations—these are optimized for speed and memory efficiency, making them perfect for big datasets:

# Convert DataFrame to numpy array and create non-NaN mask
arr = df.values
mask = df.notna().values

# Get indices that sort non-NaNs to the left (stable sort preserves original order)
sorted_indices = np.argsort(~mask, axis=1, kind='mergesort')

# Reorder the array using these indices
result_arr = arr[np.arange(len(arr))[:, None], sorted_indices]

# Convert back to DataFrame with original columns
df_result = pd.DataFrame(result_arr, columns=df.columns)

Why This Works:

  • ~mask creates a boolean array where True marks NaNs. When sorted, these True values get pushed to the right.
  • kind='mergesort' ensures a stable sort, so non-NaN values keep their original relative order (no shuffling of valid data).
  • Vectorized operations avoid looping over rows, making this approach orders of magnitude faster than row-wise apply for large datasets.

Naive (Slow) Alternative for Reference

If you're wondering why your initial apply attempt might have struggled, here's a common naive approach that works but isn't scalable:

def shift_left(row):
    non_nan = row.dropna().values
    return pd.Series(np.pad(non_nan, (0, len(row)-len(non_nan)), mode='constant', constant_values=np.nan))

# Works but is slow for 100k rows
df_result_slow = df.apply(shift_left, axis=1)

This method iterates over every row individually, which is inefficient for big data. Stick to the vectorized numpy approach for performance.

Performance Note

For a 100k-row dataset with 4 columns, the vectorized method finishes in milliseconds, while the apply method takes several seconds (or longer depending on your hardware). This makes the numpy solution ideal for production-scale data.

内容的提问来源于stack exchange,提问作者nOObda

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 10:07:15