You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python Pandas:DataFrame列非空值筛选及填充问题求助

Solution to Restructure Your DataFrame

Let's break down how to achieve your desired output in a clean, Pythonic way (no loops/conditionals required).

First, Let's Recap Your Data

Your original DataFrame s1 looks like this:

import pandas as pd
import numpy as np

s1 = pd.DataFrame([1,'a',np.nan,np.nan,np.nan,2,'b',np.nan,np.nan,np.nan,3,'c',np.nan,np.nan,np.nan]).T
# Output:
#    0  1    2    3    4
# 0  1  a  NaN  NaN  NaN
# 1  2  b  NaN  NaN  NaN
# 2  3  c  NaN  NaN  NaN

Your goal is to split each row's non-null values (numbers and letters) into separate rows, keeping only the target value in column 1 and setting all other columns to NaN.


Method 1: Directly Construct the Target DataFrame (Most Efficient)

Instead of starting with the shifted/concatenated s DataFrame, we can build the desired output directly by extracting and reshaping the non-null columns:

# Create rows for numeric values (only column 1 has data)
numeric_rows = pd.DataFrame(np.nan, index=s1.index, columns=s1.columns)
numeric_rows[1] = s1[0]

# Create rows for character values (only column 1 has data)
char_rows = pd.DataFrame(np.nan, index=s1.index, columns=s1.columns)
char_rows[1] = s1[1]

# Combine the two sets of rows
target_df = pd.concat([numeric_rows, char_rows], ignore_index=True)

This gives you exactly your desired output:

0  1    2    3    4
0  NaN  1  NaN  NaN  NaN
1  NaN  2  NaN  NaN  NaN
2  NaN  3  NaN  NaN  NaN
3  NaN  a  NaN  NaN  NaN
4  NaN  b  NaN  NaN  NaN
5  NaN  c  NaN  NaN  NaN

Method 2: Process Your Existing Concatenated DataFrame s

If you want to work with the s DataFrame you already created (from shifting and concatenating), you can use a boolean mask to keep only the values you need and set everything else to NaN:

# Your existing code to create s
s2 = s1.shift(periods=1, axis=1)
s = pd.concat([s2, s1], axis='index', join='inner', ignore_index=True, copy=False)

# Create a boolean mask to define which values to keep
mask = np.zeros_like(s, dtype=bool)
# Keep column 1 for all rows
mask[:, 1] = True
# Exclude the 'a/b/c' values in column 2 of the first 3 rows
mask[:3, 2] = False
# Exclude the '1/2/3' values in column 0 of the last 3 rows
mask[3:, 0] = False

# Apply the mask to filter the DataFrame
processed_s = s.where(mask)

This will transform your s DataFrame into the target output, ready for your subsequent ffill operation:

processed_s.fillna(method='ffill', axis=1, inplace=True)

Key Notes

Both methods use vectorized operations (no for loops or if/else statements) which are the Pythonic way to work with Pandas—they're faster and more readable than iterative approaches.

内容的提问来源于stack exchange,提问作者Yan Tian

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 07:36:09