Python Pandas:DataFrame列非空值筛选及填充问题求助
Let's break down how to achieve your desired output in a clean, Pythonic way (no loops/conditionals required).
First, Let's Recap Your Data
Your original DataFrame s1 looks like this:
import pandas as pd import numpy as np s1 = pd.DataFrame([1,'a',np.nan,np.nan,np.nan,2,'b',np.nan,np.nan,np.nan,3,'c',np.nan,np.nan,np.nan]).T # Output: # 0 1 2 3 4 # 0 1 a NaN NaN NaN # 1 2 b NaN NaN NaN # 2 3 c NaN NaN NaN
Your goal is to split each row's non-null values (numbers and letters) into separate rows, keeping only the target value in column 1 and setting all other columns to NaN.
Method 1: Directly Construct the Target DataFrame (Most Efficient)
Instead of starting with the shifted/concatenated s DataFrame, we can build the desired output directly by extracting and reshaping the non-null columns:
# Create rows for numeric values (only column 1 has data) numeric_rows = pd.DataFrame(np.nan, index=s1.index, columns=s1.columns) numeric_rows[1] = s1[0] # Create rows for character values (only column 1 has data) char_rows = pd.DataFrame(np.nan, index=s1.index, columns=s1.columns) char_rows[1] = s1[1] # Combine the two sets of rows target_df = pd.concat([numeric_rows, char_rows], ignore_index=True)
This gives you exactly your desired output:
0 1 2 3 4 0 NaN 1 NaN NaN NaN 1 NaN 2 NaN NaN NaN 2 NaN 3 NaN NaN NaN 3 NaN a NaN NaN NaN 4 NaN b NaN NaN NaN 5 NaN c NaN NaN NaN
Method 2: Process Your Existing Concatenated DataFrame s
If you want to work with the s DataFrame you already created (from shifting and concatenating), you can use a boolean mask to keep only the values you need and set everything else to NaN:
# Your existing code to create s s2 = s1.shift(periods=1, axis=1) s = pd.concat([s2, s1], axis='index', join='inner', ignore_index=True, copy=False) # Create a boolean mask to define which values to keep mask = np.zeros_like(s, dtype=bool) # Keep column 1 for all rows mask[:, 1] = True # Exclude the 'a/b/c' values in column 2 of the first 3 rows mask[:3, 2] = False # Exclude the '1/2/3' values in column 0 of the last 3 rows mask[3:, 0] = False # Apply the mask to filter the DataFrame processed_s = s.where(mask)
This will transform your s DataFrame into the target output, ready for your subsequent ffill operation:
processed_s.fillna(method='ffill', axis=1, inplace=True)
Key Notes
Both methods use vectorized operations (no for loops or if/else statements) which are the Pythonic way to work with Pandas—they're faster and more readable than iterative approaches.
内容的提问来源于stack exchange,提问作者Yan Tian

