pyspark.pandas:将float64列转为TimedeltaIndex报错求助
问题解决:pyspark.pandas转TimedeltaIndex时的KeyError错误
错误原因
你调用to_numpy()将ps.to_timedelta生成的TimedeltaSeries转换成了numpy原生的timedelta64数组,而pyspark.pandas的set_index方法无法直接识别这种numpy类型的时间差对象,因此抛出KeyError。
修复代码
直接使用ps.to_timedelta生成的TimedeltaSeries作为索引,无需转成numpy数组:
import pyspark.pandas as ps df = ps.DataFrame({"time": [2.0, 3.0, 4.0], "x": [4.5, 4.0, 3.5]}) # 直接用TimedeltaSeries设置索引,自动转为TimedeltaIndex df.set_index(ps.to_timedelta(df.time, unit="s"), inplace=True) # 或者生成新的DataFrame(不修改原数据) # new_df = df.set_index(ps.to_timedelta(df.time, unit="s"))
验证索引类型
设置完成后可以通过以下代码确认索引类型是ps.TimedeltaIndex,满足重采样需求:
print(type(df.index)) # 输出:<class 'pyspark.pandas.indexes.timedeltas.TimedeltaIndex'>
内容的提问来源于stack exchange,提问作者ascripter
相关产品推荐
相关产品推荐

