如何返回DataFrame每行的最早非NaN日期及多值行索引?
Got it, let's work through this problem to get exactly what you need—keeping the all-NaN row, grabbing the correct date for each row, and identifying rows with multiple non-NaN values.
Step 1: Prepare the DataFrame
First, let's convert all date columns to datetime type so we can properly compare and find the earliest date:
import pandas as pd import numpy as np data = [[np.nan, np.nan, np.nan], [np.nan, '1987-09-01', np.nan], [np.nan, '1987-09-01', '1988-09-01']] df = pd.DataFrame(data, columns=['1', '2', '3']) # Convert all columns to datetime to enable proper date comparisons df = df.apply(pd.to_datetime)
Step 2: Get the Earliest Date per Row (Including All-NaN Rows)
The min() method on rows works perfectly here—it returns the earliest date for rows with values, and NaT (datetime equivalent of NaN) for all-NaN rows, which is exactly what you want:
# Calculate earliest date per row; preserves all rows including all-NaN ones earliest_dates = df.min(axis=1)
Running this gives you:
0 NaT 1 1987-09-01 2 1987-09-01 dtype: datetime64[ns]
Unlike stack().groupby(level=0).first(), this doesn't drop the all-NaN row, and it reliably picks the earliest date (not just the first one encountered in column order).
Step 3: Identify Rows with Multiple Non-NaN Values
To find which rows have more than one valid date, we can count non-null values per row and filter for counts ≥ 2:
# Count non-null values in each row non_null_counts = df.notna().sum(axis=1) # Extract indices of rows with 2+ non-null values multi_value_indices = non_null_counts[non_null_counts >= 2].index.tolist()
This will return [2] for your sample data, which is the index of the row with two dates.
Putting It All Together
If you want to inspect the results together:
print("Earliest date for each row:") print(earliest_dates) print("\nIndices of rows with multiple values:") print(multi_value_indices)
Why Your Original Approach Didn't Work
df.stack() defaults to dropping rows with all NaN values (even if you add dropna=False, groupby(...).first() would just return the first non-null value in column order—not necessarily the earliest date). Using min(axis=1) is more straightforward and aligned with your "earliest date" requirement.
内容的提问来源于stack exchange,提问作者bjornvandijkman

