如何解决resample后Series使用apply()多聚合操作的报错?
解决Resample后Apply报错'Series' object has no attribute 'columns'
问题背景
对DataFrame执行重采样后使用apply()方法时,因操作对象变为Series而非DataFrame,触发如下错误:
'Series' object has no attribute 'columns'
报错代码
df = pd.DataFrame(data) x = df[df.columns[0]].fillna(0) x_r = x.resample("s") x = x_r.apply(['mean', np.max, np.min])
完整报错栈
--------------------------------------------------------------------------- AttributeError Traceback (most recent call last) /tmp/ipykernel_296/635968.py in <module> 2 print(x_mm) 3 ----> 4 x_mm.apply(['mean',np.max,np.min]) /local/opt/anaconda/anaconda3/envs/lib/python3.7/site-packages/pandas/core/resample.py in aggregate(self, func, *args, **kwargs) 333 def aggregate(self, func, *args, **kwargs): 334 ---> 335 result = ResamplerWindowApply(self, func, args=args, kwargs=kwargs).agg() 336 if result is None: 337 how = func /local/opt/anaconda/anaconda3/envs/lib/python3.7/site-packages/pandas/core/apply.py in agg(self) 162 elif is_list_like(arg): 163 # we require a list, but not a 'str' ---> 164 return self.agg_list_like() 165 166 if callable(arg): /local/opt/anaconda/anaconda3/envs/lib/python3.7/site-packages/pandas/core/apply.py in agg_list_like(self) 334 if selected_obj.ndim == 1: 335 for a in arg: ---> 336 colg = obj._gotitem(selected_obj.name, ndim=1, subset=selected_obj) 337 try: 338 new_res = colg.aggregate(a) /local/opt/anaconda/anaconda3/envs/lib/python3.7/site-packages/pandas/core/resample.py in _gotitem(self, key, ndim, subset) 390 # try the key selection 391 try: ---> 392 return grouped[key] 393 except KeyError: 394 return grouped /local/opt/anaconda/anaconda3/envs/lib/python3.7/site-packages/pandas/core/base.py in __getitem__(self, key) 218 219 if isinstance(key, (list, tuple, ABCSeries, ABCIndex, np.ndarray)): ---> 220 if len(self.obj.columns.intersection(key)) != len(key): 221 bad_keys = list(set(key).difference(self.obj.columns)) 222 raise KeyError(f"Columns not found: {str(bad_keys)[1:-1]}") /local/opt/anaconda/anaconda3/envs/lib/python3.7/site-packages/pandas/core/generic.py in __getattr__(self, name) 5485 ): 5486 return self[name] -> 5487 return object.__getattribute__(self, name) 5488 5489 def __setattr__(self, name: str, value) -> None: AttributeError: 'Series' object has no attribute 'columns'
原因分析
代码中df[df.columns[0]]取DataFrame单列时返回的是Series对象,旧版本pandas中,对Series的Resampler传入列表形式的聚合函数(如['mean', np.max, np.min])时,内部逻辑会尝试访问columns属性,而Series没有该属性,因此报错。
解决方案
方案1:保留DataFrame结构
使用双层方括号选取单列,确保操作对象始终是DataFrame:
import pandas as pd import numpy as np df = pd.DataFrame(data) # 双层方括号保留DataFrame结构,而非转为Series x = df[[df.columns[0]]].fillna(0) x_r = x.resample("s") x = x_r.apply(['mean', np.max, np.min])
方案2:使用agg方法替代apply(推荐)
对于聚合操作,agg()比apply()更适配列表形式的聚合函数,且对Series同样友好:
import pandas as pd import numpy as np df = pd.DataFrame(data) x = df[df.columns[0]].fillna(0) x_r = x.resample("s") # 用agg直接传入聚合函数列表 x = x_r.agg(['mean', np.max, np.min])
方案3:自定义函数处理Series
如果必须使用apply(),可以自定义函数返回包含多个统计量的Series:
import pandas as pd import numpy as np def calculate_stats(s): return pd.Series({ 'mean': s.mean(), 'max': np.max(s), 'min': np.min(s) }) df = pd.DataFrame(data) x = df[df.columns[0]].fillna(0) x_r = x.resample("s") x = x_r.apply(calculate_stats)
内容的提问来源于stack exchange,提问作者Art
相关产品推荐
相关产品推荐

