You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何返回DataFrame每行的最早非NaN日期及多值行索引?

Solution for Your Pandas Date Handling Problem

Got it, let's work through this problem to get exactly what you need—keeping the all-NaN row, grabbing the correct date for each row, and identifying rows with multiple non-NaN values.

Step 1: Prepare the DataFrame

First, let's convert all date columns to datetime type so we can properly compare and find the earliest date:

import pandas as pd
import numpy as np

data = [[np.nan, np.nan, np.nan], [np.nan, '1987-09-01', np.nan], [np.nan, '1987-09-01', '1988-09-01']]
df = pd.DataFrame(data, columns=['1', '2', '3'])

# Convert all columns to datetime to enable proper date comparisons
df = df.apply(pd.to_datetime)

Step 2: Get the Earliest Date per Row (Including All-NaN Rows)

The min() method on rows works perfectly here—it returns the earliest date for rows with values, and NaT (datetime equivalent of NaN) for all-NaN rows, which is exactly what you want:

# Calculate earliest date per row; preserves all rows including all-NaN ones
earliest_dates = df.min(axis=1)

Running this gives you:

0          NaT
1   1987-09-01
2   1987-09-01
dtype: datetime64[ns]

Unlike stack().groupby(level=0).first(), this doesn't drop the all-NaN row, and it reliably picks the earliest date (not just the first one encountered in column order).

Step 3: Identify Rows with Multiple Non-NaN Values

To find which rows have more than one valid date, we can count non-null values per row and filter for counts ≥ 2:

# Count non-null values in each row
non_null_counts = df.notna().sum(axis=1)

# Extract indices of rows with 2+ non-null values
multi_value_indices = non_null_counts[non_null_counts >= 2].index.tolist()

This will return [2] for your sample data, which is the index of the row with two dates.

Putting It All Together

If you want to inspect the results together:

print("Earliest date for each row:")
print(earliest_dates)
print("\nIndices of rows with multiple values:")
print(multi_value_indices)

Why Your Original Approach Didn't Work

df.stack() defaults to dropping rows with all NaN values (even if you add dropna=False, groupby(...).first() would just return the first non-null value in column order—not necessarily the earliest date). Using min(axis=1) is more straightforward and aligned with your "earliest date" requirement.

内容的提问来源于stack exchange,提问作者bjornvandijkman

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 08:50:36