Pandas代码报错:timedelta64[ns]与int无效比较的解决方法
问题描述
运行以下Python Pandas代码时触发TypeError:
for file in files: df=pd.read_csv(file) df["time_stamp"]=pd.to_datetime(df["time_stamp"], infer_datetime_format=True) df= df.resample('1Min', on='time_stamp').first().dropna() df["time_diff"]=df["time_stamp"].diff() df=df[1:] df["time_diff_h"]=df["time_diff"].apply(lambda x: x.total_seconds()/3600) df=df.reset_index(drop=True) #index_list=df[df["time_diff_h"]>time_gap].index index_list = df[df["time_diff_h"]>24].index if len(index_list)==0: continue df_output=pd.DataFrame() for i in index_list: soc_diff=df["ibs3_soc__0x7a3 (%)"].iloc[i]-df["ibs3_soc__0x7a3 (%)"].iloc[i-1] initial_soc=df["ibs3_soc__0x7a3 (%)"].iloc[i-1] final_soc=df["ibs3_soc__0x7a3 (%)"].iloc[i] timeDiff= round((df["time_diff_h"].iloc[i]),2) initial_time=df["time_stamp"].iloc[i-1] final_time=df["time_stamp"].iloc[i] power=(df["ibs3_q_released__0x7a3 (ah)"].iloc[i]-df["ibs3_q_released__0x7a3 (ah)"].iloc[i-1])/timeDiff if power<0.125: continue else: features={"File": file, "SOC_Difference (%)": soc_diff, "Initial_SOC (%)": initial_soc, "Final_SOC (%)": final_soc, "Time_Diff (h)": timeDiff, "Initial_time": initial_time, "Final_time": final_time, "Power_released (Ah)": power} df_output=df_output.append(pd.Series(features), ignore_index=True)
报错信息:
raise TypeError(f"Invalid comparison between dtype={left.dtype} and {typ}") TypeError: Invalid comparison between dtype=timedelta64[ns] and int
已通过代码排查确认仅time_diff列为timedelta64[ns]类型,且未直接对该列执行数值比较,无法定位报错原因,寻求解决方案。
解决方法
1. 替换apply为矢量化时间差转换
apply属于逐元素操作,容易因个别异常值导致列类型混乱。改用Pandas内置的dt访问器做矢量化转换,确保输出为数值类型:
# 替换原time_diff_h生成代码 df["time_diff_h"] = df["time_diff"].dt.total_seconds() / 3600
2. 修正resample后的索引处理
resample操作后time_stamp会成为索引,后续直接调用df["time_stamp"].diff()可能触发隐式类型转换异常。调整resample代码,显式重置索引:
df = df.resample('1Min', on='time_stamp').first().dropna().reset_index()
3. 强制比较两边类型一致
在生成index_list时,将整数24显式转为浮点型,避免类型匹配问题:
index_list = df[df["time_diff_h"]>24.0].index
4. 临时排查:打印关键列类型
在报错代码前添加打印语句,确认time_diff_h的实际类型和样本值,快速定位异常:
print("time_diff_h dtype:", df['time_diff_h'].dtype) print("time_diff_h sample:", df['time_diff_h'].head())
内容的提问来源于stack exchange,提问作者AC24
相关产品推荐
相关产品推荐

