使用pandas get_loc(ffill)查询早于起始日期触发KeyError问题
get_loc('ffill') Hey there! Let's break down why you're hitting this KeyError and how to adjust your code to return None instead when the input date is earlier than your DataFrame's first entry.
Why the KeyError Happens
When you use get_loc(dt, 'ffill'), pandas tries to find the latest index entry that's less than or equal to your target date (that's what "forward fill" means here). But if your target date is earlier than the very first entry in your index, there's no such entry to fall back on—pandas can't "fill forward" from nothing, so it throws a KeyError.
Dates later than the last entry work because ffill just uses the final index entry, but there's no equivalent default for dates before the first entry.
Solution 1: Add a Pre-Check for Boundary Dates
The simplest fix is to explicitly check if your target date is before the index's minimum value. If it is, return None right away; otherwise, proceed with get_loc:
df_colname = 'Date Time' pandas_datetime_colname = 'Pandas Date Time' df[pandas_datetime_colname] = pd.to_datetime(df[df_colname]) df.set_index(pandas_datetime_colname, inplace=True) dt = pd.to_datetime(inputdatetime) # Check if date is before the first index entry if dt < df.index.min(): idx = None else: idx = df.index.get_loc(dt, 'ffill') print(f"Date Time: {inputdatetime} :idx {idx}") df.reset_index(inplace=True)
Solution 2: Use get_indexer for Gentler Boundary Handling
Another approach is to use pd.Index.get_indexer() instead of get_loc(). This method returns -1 for dates that fall outside the index bounds (including before the first entry) instead of throwing an error. You can then convert -1 to None:
df_colname = 'Date Time' pandas_datetime_colname = 'Pandas Date Time' df[pandas_datetime_colname] = pd.to_datetime(df[df_colname]) df.set_index(pandas_datetime_colname, inplace=True) dt = pd.to_datetime(inputdatetime) # get_indexer returns an array, so we grab the first element idx = df.index.get_indexer([dt], method='ffill')[0] # Convert -1 (out of bounds) to None idx = idx if idx != -1 else None print(f"Date Time: {inputdatetime} :idx {idx}") df.reset_index(inplace=True)
This method is cleaner if you want to handle both lower and upper bounds consistently (though your upper bound already works as expected).
Testing It Out
- For a date in your data range (e.g.,
2020-02-02 22:50:00), both solutions return the correct index. - For a date later than your last entry,
ffillwill return the index of the final row (which matches your existing behavior). - For a date earlier than your first entry (e.g.,
2019-12-20 22:45:00), both solutions returnNoneinstead of throwing a KeyError.
内容的提问来源于stack exchange,提问作者Ramana

