如何实现类似Series.asof但取首个无NaN值行的功能?
Great question! Since Series.asof() pulls the last non-null value up to a given index, getting the first one depends a bit on exactly what you need. Let's break down two common scenarios and how to solve them:
1. Get the very first non-null value in the entire Series
If you just need the earliest non-null entry across your whole Series, pandas has a straightforward built-in way to do this with first_valid_index():
import pandas as pd import numpy as np # Example Series s = pd.Series([np.nan, 10, np.nan, 20, np.nan], index=[0, 1, 2, 3, 4]) # Grab the index of the first non-null value first_valid_idx = s.first_valid_index() # Get the corresponding value first_valid_value = s.loc[first_valid_idx] print(first_valid_value) # Output: 10
If your Series has no non-null values at all, first_valid_index() will return None, so you might want to add a quick check for that edge case if your data could have it.
2. Mimic asof() behavior for batch queries (get first non-null up to each query index)
If you need to replicate asof()'s batch query functionality—meaning you pass multiple index values and get a result for each—you have two common sub-scenarios to consider:
Scenario A: Get the earliest non-null value ever up to each query index
This returns the first non-null value in the entire Series for any query index that comes after that first non-null entry. For example, if your first non-null is at index 1, every query index >=1 will return that value.
Here's a simple implementation:
def first_asof_global(s, where): # Extract all non-null entries (sorted by index) valid_entries = s.dropna() if valid_entries.empty: return pd.Series([np.nan]*len(where), index=where) first_valid_val = valid_entries.iloc[0] first_valid_idx = valid_entries.index[0] # For each query index, return the first valid value if query >= first valid index, else NaN results = [first_valid_val if idx >= first_valid_idx else np.nan for idx in where] return pd.Series(results, index=where) # Test it s = pd.Series([np.nan, 10, np.nan, 20, np.nan, 30], index=[0,1,2,3,4,5]) print(first_asof_global(s, where=[2, 3, 5])) # Output: 10, 10, 10
Scenario B: Get the first value of the last contiguous non-null block up to each query index
This is useful if your Series has chunks of non-null values separated by NaNs. Instead of returning the very first non-null ever, it returns the start of the most recent non-null block that ends at or before your query index. This is the closest parallel to asof() (which returns the end of that block).
Here's how to implement it:
def first_asof_contiguous(s, where): # Create a mask of non-null values non_null_mask = s.notna() # Identify the start of each contiguous non-null block block_starts = non_null_mask & ~non_null_mask.shift(fill_value=False) # Extract the values at the start of each block block_start_values = s[block_starts] if block_start_values.empty: return pd.Series([np.nan]*len(where), index=where) # Use asof() logic on the block start values to find the last block start <= each query index return block_start_values.asof(where) # Test it s = pd.Series([np.nan, 10, 15, np.nan, 20, 25, np.nan], index=[0,1,2,3,4,5,6]) print(first_asof_contiguous(s, where=[2, 4, 5])) # Output: 10, 20, 20 # Compare to s.asof([2,4,5]) which returns: 15, 25, 25
Pick the approach that matches your specific use case—either grabbing the global first non-null, or the start of the most recent non-null block before each query point.
内容的提问来源于stack exchange,提问作者Stephen Zhou

