You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas中DatetimeIndex调用get_loc报InvalidIndexError问题排查

问题描述

我有一个来自Pandas DataFrame索引的源时间戳列表,想要在另外两个目标DataFrame中查找最近的时间戳。其中target2调用index.get_loc(ts, "nearest")能正常运行,但target1却返回错误:InvalidIndexError: Reindexing only valid with uniquely valued Index objects。原本以为已对两个目标索引排序并去重,后续发现对应target1的table1存在重复索引,想了解问题根源。

相关代码及输出如下:

source = df.set_index(['timestamp']).sort_index().index
source[10000:10005]
# 输出:
DatetimeIndex(['2023-04-21 19:35:59.396250', '2023-04-21 19:35:59.475000',
               '2023-04-21 19:35:59.555000', '2023-04-21 19:35:59.635000',
               '2023-04-21 19:35:59.716250'],
              dtype='datetime64[ns]', name='timestamp', freq=None)

# 尝试对两个目标DataFrame排序并去重
target1 = df1.set_index(['timestamp']).sort_index().drop_duplicates()
target2 = df2.set_index(['timestamp']).sort_index().drop_duplicates()

target1.index[5:]
# 输出:
DatetimeIndex(['2023-04-21 19:27:30.869530', '2023-04-21 19:27:31.116220',
               '2023-04-21 19:27:31.364920', '2023-04-21 19:27:31.616380',
               '2023-04-21 19:27:31.880470'],
              dtype='datetime64[ns]', name='timestamp', freq=None)
target1.index.max()
# 输出:Timestamp('2023-04-21 19:58:50.872560')

target2.index[5:]
# 输出:
DatetimeIndex(['2023-04-21 18:58:56.587016', '2023-04-21 18:58:57.587014',
               '2023-04-21 18:58:58.587013', '2023-04-21 18:58:59.587014',
               '2023-04-21 18:59:00.587012'],
              dtype='datetime64[ns]', name='timestamp', freq=None)
target2.index.max()
# 输出:Timestamp('2023-04-21 19:56:48.581202')

循环测试结果:

# target2测试正常
for ts in source[10000:10005]:
    get = target2.index.get_loc(ts,"nearest")
    print("Source:",ts, "Nearest:",target2[get])

# 输出:
Source: 2023-04-21 19:35:59.396250 Nearest: 2023-04-21 19:35:59.581398
Source: 2023-04-21 19:35:59.475000 Nearest: 2023-04-21 19:35:59.581398
Source: 2023-04-21 19:35:59.555000 Nearest: 2023-04-21 19:35:59.581398
Source: 2023-04-21 19:35:59.635000 Nearest: 2023-04-21 19:35:59.581398
Source: 2023-04-21 19:35:59.716250 Nearest: 2023-04-21 19:35:59.581398

# target1测试报错
for ts in source[10000:10005]:
    get = target1.index.get_loc(ts,"nearest")
    print("Source:",ts, "Nearest:",target1[get])

# 报错信息:
---------------------------------------------------------------------------
InvalidIndexError                         Traceback (most recent call last)
~\AppData\Local\Temp/ipykernel_19096/211410640.py in <module>
      1 for ts in source[10000:10005]:
----> 2     get = target1.get_loc(t,"nearest")
      3     print("Source:",ts, "Nearest:",target1[get])

~\Anaconda3\lib\site-packages\pandas\core\indexes\datetimes.py in get_loc(self, key, method, tolerance)
    701 
    702         try:
---> 703             return Index.get_loc(self, key, method, tolerance)
    704         except KeyError as err:
    705             raise KeyError(orig_key) from err

~\Anaconda3\lib\site-packages\pandas\core\indexes\base.py in get_loc(self, key, method, tolerance)
   3369             tolerance = self._convert_tolerance(tolerance, np.asarray(key))
   3370 
-> 3371         indexer = self.get_indexer([key], method=method, tolerance=tolerance)
   3372         if indexer.ndim > 1 or indexer.size > 1:
   3373             raise TypeError("get_loc requires scalar valued input")

~\Anaconda3\lib\site-packages\pandas\core\indexes\base.py in get_indexer(self, target, method, limit, tolerance)
   3440 
   3441         if not self._index_as_unique:
-> 3442             raise InvalidIndexError(self._requires_unique_msg)
   3443 
   3444         if not self._should_compare(target) and not is_interval_dtype(self.dtype):

InvalidIndexError: Reindexing only valid with uniquely valued Index objects

更新信息:table1是ping结果的DataFrame,确实存在重复索引:

table1[table1.index.duplicated()].head()
# 输出:
                     Name    Client  Target   seq  bytes   delay
timestamp                                                       
2023-04-21 19:28:52.967250  fping  test11041  1.1.1.1   328     64  144.0
2023-04-21 19:34:37.321620  fping  test11041  1.1.1.1  1703     64  749.0
2023-04-21 19:34:37.321620  fping  test11041  1.1.1.1  1705     64  249.0
2023-04-21 19:35:48.729640  fping  test11041  1.1.1.1  1991     64  157.0
2023-04-21 19:36:52.478800  fping  test11041  1.1.1.1  2246     64  156.0
问题根源与解决方法

问题根源

  1. drop_duplicates()的作用对象错误:你调用的DataFrame.drop_duplicates()是基于整行所有列的数据判断重复,而非仅索引。当table1中存在相同时间戳但其他列(如seq、delay)值不同的行时,drop_duplicates()不会将这些行判定为重复行,因此target1的索引仍会保留重复值。
  2. get_loc(method="nearest")的要求:Pandas的这个方法要求索引必须是唯一的。如果索引存在重复,无法确定返回哪个重复位置的“最近”值,因此会抛出InvalidIndexError。

解决方法

方法一:直接对索引去重

针对索引本身进行去重操作,保留每个时间戳对应的第一行(或最后一行):

target1 = df1.set_index(['timestamp']).sort_index()
# 对索引去重,keep='first'保留第一个出现的行,'last'保留最后一个
target1 = target1[~target1.index.duplicated(keep='first')]

方法二:设置索引前先对时间戳列去重

先基于timestamp列去重,再设置索引,确保索引唯一:

# subset指定只根据timestamp列判断重复
target1 = df1.drop_duplicates(subset=['timestamp']).set_index(['timestamp']).sort_index()

验证索引唯一性

处理后可以通过以下代码确认索引是否唯一:

print(target1.index.is_unique)  # 返回True说明索引已无重复

内容的提问来源于stack exchange,提问作者senor_smiley

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.23 18:44:55