You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用pandas get_loc(ffill)查询早于起始日期触发KeyError问题

Fixing KeyError for Dates Earlier Than the First Index Entry with get_loc('ffill')

Hey there! Let's break down why you're hitting this KeyError and how to adjust your code to return None instead when the input date is earlier than your DataFrame's first entry.

Why the KeyError Happens

When you use get_loc(dt, 'ffill'), pandas tries to find the latest index entry that's less than or equal to your target date (that's what "forward fill" means here). But if your target date is earlier than the very first entry in your index, there's no such entry to fall back on—pandas can't "fill forward" from nothing, so it throws a KeyError.

Dates later than the last entry work because ffill just uses the final index entry, but there's no equivalent default for dates before the first entry.

Solution 1: Add a Pre-Check for Boundary Dates

The simplest fix is to explicitly check if your target date is before the index's minimum value. If it is, return None right away; otherwise, proceed with get_loc:

df_colname = 'Date Time'
pandas_datetime_colname = 'Pandas Date Time'
df[pandas_datetime_colname] = pd.to_datetime(df[df_colname])
df.set_index(pandas_datetime_colname, inplace=True)

dt = pd.to_datetime(inputdatetime)
# Check if date is before the first index entry
if dt < df.index.min():
    idx = None
else:
    idx = df.index.get_loc(dt, 'ffill')

print(f"Date Time: {inputdatetime} :idx {idx}")
df.reset_index(inplace=True)

Solution 2: Use get_indexer for Gentler Boundary Handling

Another approach is to use pd.Index.get_indexer() instead of get_loc(). This method returns -1 for dates that fall outside the index bounds (including before the first entry) instead of throwing an error. You can then convert -1 to None:

df_colname = 'Date Time'
pandas_datetime_colname = 'Pandas Date Time'
df[pandas_datetime_colname] = pd.to_datetime(df[df_colname])
df.set_index(pandas_datetime_colname, inplace=True)

dt = pd.to_datetime(inputdatetime)
# get_indexer returns an array, so we grab the first element
idx = df.index.get_indexer([dt], method='ffill')[0]
# Convert -1 (out of bounds) to None
idx = idx if idx != -1 else None

print(f"Date Time: {inputdatetime} :idx {idx}")
df.reset_index(inplace=True)

This method is cleaner if you want to handle both lower and upper bounds consistently (though your upper bound already works as expected).

Testing It Out

  • For a date in your data range (e.g., 2020-02-02 22:50:00), both solutions return the correct index.
  • For a date later than your last entry, ffill will return the index of the final row (which matches your existing behavior).
  • For a date earlier than your first entry (e.g., 2019-12-20 22:45:00), both solutions return None instead of throwing a KeyError.

内容的提问来源于stack exchange,提问作者Ramana

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.06 15:02:56