如何解决TypeError: 'last'仅支持DatetimeIndex索引错误
问题排查:pandas last('3M')触发TypeError错误
运行日志解析代码时触发TypeError,错误提示:'last' only supports a DatetimeIndex index,目标是用last('3M')筛选最近3个月数据但执行失败。
原代码
def create_excel_file(): master_list = [] for name in filelist: new_path = Path(name).parent base = os.path.basename(new_path) final = os.path.splitext(base)[0] with open(name,"r") as f: soupObj = bs4.BeautifulSoup(f, "lxml") df = pd.DataFrame([(x["uri"], *x["t"].split("T"), x["u"], x["desc"]) for x in soupObj.find_all("log")], columns=["Document", "Date", "Time", "User", "Description"]) df.insert(0, 'Database', f'{final}') df['Document'] = df['Document'].astype(str) df['Date'] = pd.to_datetime(df['Date']).dt.date master_list.append(df) df = pd.concat(master_list, axis=0, ignore_index=True) df = df.sort_values(by='Date', ascending=True).set_index('Date').last('3M') df = df.sort_values(by='Date', ascending=False) df.to_excel("logfile.xlsx", index=True) create_excel_file()
报错信息
Traceback (most recent call last): File "C:\Users\Desktop\project\Final test.py", line 40, in <module> create_excel_file() File "C:\Users\Desktop\project\Final test.py", line 34, in create_excel_file df = df.sort_values(by='Date', ascending=True).set_index('Date').last('3M') ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "C:\Users\AppData\Roaming\Python\Python311\site-packages\pandas\core\generic.py", line 9001, in last raise TypeError("'last' only supports a DatetimeIndex index") TypeError: 'last' only supports a DatetimeIndex index Process finished with exit code 1
问题原因
代码里把Date列转成了Python原生的date对象(通过pd.to_datetime(df['Date']).dt.date),将其设置为索引后,索引类型是ObjectIndex而非pandas要求的DatetimeIndex,而last()方法仅支持DatetimeIndex类型的索引,因此触发错误。
解决办法
去掉dt.date转换,让Date列保持pandas的datetime64类型,这样设置为索引后就是DatetimeIndex,就能正常使用last('3M')了。
修改后的完整代码:
def create_excel_file(): master_list = [] for name in filelist: new_path = Path(name).parent base = os.path.basename(new_path) final = os.path.splitext(base)[0] with open(name,"r") as f: soupObj = bs4.BeautifulSoup(f, "lxml") df = pd.DataFrame([(x["uri"], *x["t"].split("T"), x["u"], x["desc"]) for x in soupObj.find_all("log")], columns=["Document", "Date", "Time", "User", "Description"]) df.insert(0, 'Database', f'{final}') df['Document'] = df['Document'].astype(str) # 去掉.dt.date,保留datetime类型 df['Date'] = pd.to_datetime(df['Date']) master_list.append(df) df = pd.concat(master_list, axis=0, ignore_index=True) df = df.sort_values(by='Date', ascending=True).set_index('Date').last('3M') df = df.sort_values(by='Date', ascending=False) df.to_excel("logfile.xlsx", index=True) create_excel_file()
如果需要最终Excel里的日期显示为纯日期格式(不带时间),可以在写入Excel时设置格式,或者在后续处理中用dt.date但不要把它设为索引,改用布尔索引筛选最近3个月数据:
示例(替代last('3M')的写法):
# 保留dt.date转换,但不设为索引,用布尔筛选 df['Date'] = pd.to_datetime(df['Date']).dt.date df = pd.concat(master_list, axis=0, ignore_index=True) # 计算3个月前的日期 three_months_ago = pd.Timestamp.now() - pd.DateOffset(months=3) # 转换为date对象匹配列类型 three_months_ago = three_months_ago.date() # 筛选最近3个月数据 df = df[df['Date'] >= three_months_ago].sort_values(by='Date', ascending=False)
内容的提问来源于stack exchange,提问作者Rejoy
相关产品推荐
相关产品推荐

