You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何实现类似Series.asof但取首个无NaN值行的功能?

Great question! Since Series.asof() pulls the last non-null value up to a given index, getting the first one depends a bit on exactly what you need. Let's break down two common scenarios and how to solve them:

1. Get the very first non-null value in the entire Series

If you just need the earliest non-null entry across your whole Series, pandas has a straightforward built-in way to do this with first_valid_index():

import pandas as pd
import numpy as np

# Example Series
s = pd.Series([np.nan, 10, np.nan, 20, np.nan], index=[0, 1, 2, 3, 4])

# Grab the index of the first non-null value
first_valid_idx = s.first_valid_index()
# Get the corresponding value
first_valid_value = s.loc[first_valid_idx]

print(first_valid_value)  # Output: 10

If your Series has no non-null values at all, first_valid_index() will return None, so you might want to add a quick check for that edge case if your data could have it.

2. Mimic asof() behavior for batch queries (get first non-null up to each query index)

If you need to replicate asof()'s batch query functionality—meaning you pass multiple index values and get a result for each—you have two common sub-scenarios to consider:

Scenario A: Get the earliest non-null value ever up to each query index

This returns the first non-null value in the entire Series for any query index that comes after that first non-null entry. For example, if your first non-null is at index 1, every query index >=1 will return that value.

Here's a simple implementation:

def first_asof_global(s, where):
    # Extract all non-null entries (sorted by index)
    valid_entries = s.dropna()
    if valid_entries.empty:
        return pd.Series([np.nan]*len(where), index=where)
    
    first_valid_val = valid_entries.iloc[0]
    first_valid_idx = valid_entries.index[0]
    
    # For each query index, return the first valid value if query >= first valid index, else NaN
    results = [first_valid_val if idx >= first_valid_idx else np.nan for idx in where]
    return pd.Series(results, index=where)

# Test it
s = pd.Series([np.nan, 10, np.nan, 20, np.nan, 30], index=[0,1,2,3,4,5])
print(first_asof_global(s, where=[2, 3, 5]))  # Output: 10, 10, 10

Scenario B: Get the first value of the last contiguous non-null block up to each query index

This is useful if your Series has chunks of non-null values separated by NaNs. Instead of returning the very first non-null ever, it returns the start of the most recent non-null block that ends at or before your query index. This is the closest parallel to asof() (which returns the end of that block).

Here's how to implement it:

def first_asof_contiguous(s, where):
    # Create a mask of non-null values
    non_null_mask = s.notna()
    # Identify the start of each contiguous non-null block
    block_starts = non_null_mask & ~non_null_mask.shift(fill_value=False)
    # Extract the values at the start of each block
    block_start_values = s[block_starts]
    
    if block_start_values.empty:
        return pd.Series([np.nan]*len(where), index=where)
    
    # Use asof() logic on the block start values to find the last block start <= each query index
    return block_start_values.asof(where)

# Test it
s = pd.Series([np.nan, 10, 15, np.nan, 20, 25, np.nan], index=[0,1,2,3,4,5,6])
print(first_asof_contiguous(s, where=[2, 4, 5]))  # Output: 10, 20, 20
# Compare to s.asof([2,4,5]) which returns: 15, 25, 25

Pick the approach that matches your specific use case—either grabbing the global first non-null, or the start of the most recent non-null block before each query point.

内容的提问来源于stack exchange,提问作者Stephen Zhou

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 10:49:50