Spark上Pandas API重采样报错:小时级规则不支持
在Databricks Runtime 15.4中使用PySpark Pandas API重采样时的小时粒度规则报错问题
我在Databricks Runtime 15.4(Spark 3.5.0)环境下,尝试用Spark上的Pandas API(pyspark.pandas)执行重采样操作,测试代码如下:
import numpy as np from datetime import datetime import pandas as pd import pyspark.pandas as ps dates = [ datetime(2022, 5, 1, 4, 5, 6), datetime(2022, 5, 3), datetime(2022, 5, 3, 23, 59, 59), datetime(2022, 5, 4), pd.NaT, datetime(2022, 5, 4, 0, 0, 1), datetime(2022, 5, 11), ] df = ps.DataFrame(np.random.rand(len(dates), 2), index=pd.DatetimeIndex(dates), columns=["A", "B"]) print(df) df.resample('2H').mean()
原本是简单的重采样任务,但执行时出现了无法理解的报错:
File /databricks/spark/python/pyspark/pandas/resample.py:724, in DataFrameResampler.__init__(self, psdf, resamplekey, rule, closed, label, agg_columns) 715 def __init__( 716 self, 717 psdf: DataFrame, (...) 722 agg_columns: List[Series] = [], 723 ): --> 724 super().__init__( 725 psdf=psdf, 726 resamplekey=resamplekey, 727 rule=rule, 728 closed=closed, 729 label=label, 730 agg_columns=agg_columns, 731 ) File /databricks/spark/python/pyspark/pandas/resample.py:102, in Resampler.__init__(self, psdf, resamplekey, rule, closed, label, agg_columns) 100 self._offset = to_offset(rule) 101 if self._offset.rule_code not in ["A-DEC", "M", "D", "H", "T", "S"]: --> 102 raise ValueError("rule code {} is not supported".format(self._offset.rule_code)) 103 if not getattr(self._offset, "n") > 0: 104 raise ValueError("rule offset must be positive")
使用天(D)、月(M)等规则码时操作正常,但小时这类细粒度规则会触发报错。更换不同Databricks运行时版本后问题依旧,单独调用to_offset('2H')没有异常:
from pandas.tseries.frequencies import to_offset to_offset('2H')
内容的提问来源于stack exchange,提问作者Mariusz Jarczak
相关产品推荐
相关产品推荐

