distfit运行fit_transform报object无dtype属性错误如何解决
问题背景
我正在尝试复现《如何使用Python确定最优拟合数据分布》一文中描述的结果,使用的代码如下:
import numpy as np from distfit import distfit # 生成10000个均值为0、标准差为3的正态分布样本 X = np.random.normal(0, 3, 10000) # 初始化distfit dist = distfit() # 计算数据的最优拟合概率分布 dist.fit_transform(X)
报错信息
在Jupyter环境运行上述代码时,抛出如下错误:
[distfit] >fit.. [distfit] >transform.. --------------------------------------------------------------------------- AttributeError Traceback (most recent call last) <ipython-input-8-02f73e7f157d> in <module> 9 10 # Determine best-fitting probability distribution for data ---> 11 dist.fit_transform(X) ~\Anaconda3\lib\site-packages\distfit\distfit.py in fit_transform(self, X, verbose) 275 self.fit(verbose=verbose) 276 # Transform X based on functions ---> 277 self.transform(X, verbose=verbose) 278 # Store 279 results = _store(self.alpha, ~\Anaconda3\lib\site-packages\distfit\distfit.py in transform(self, X, verbose) 214 if self.method=='parametric': 215 # Compute best distribution fit on the empirical X ---> 216 out_summary, model = _compute_score_distribution(X, X_bins, y_obs, self.distributions, self.stats, verbose=verbose) 217 # Determine confidence intervals on the best fitting distribution 218 model = _compute_cii(self, model, verbose=verbose) ~\Anaconda3\lib\site-packages\distfit\distfit.py in _compute_score_distribution(data, X, y_obs, DISTRIBUTIONS, stats, verbose) 906 model['params'] = (0.0, 1.0) 907 best_score = np.inf ---> 908 df = pd.DataFrame(index=range(0, len(DISTRIBUTIONS)), columns=['distr', 'score', 'LLE', 'loc', 'scale', 'arg']) 909 max_name_len = np.max(list(map(lambda x: len(x.name), DISTRIBUTIONS))) 910 ~\Anaconda3\lib\site-packages\pandas\core\frame.py in __init__(self, data, index, columns, dtype, copy) 346 dtype=dtype, copy=copy) 347 elif isinstance(data, dict): ---> 348 mgr = self._init_dict(data, index, columns, dtype=dtype) 349 elif isinstance(data, ma.MaskedArray): 350 import numpy.ma.mrecords as mrecords ~\Anaconda3\lib\site-packages\pandas\core\frame.py in _init_dict(self, data, index, columns, dtype) 449 nan_dtype = dtype 450 v = construct_1d_arraylike_from_scalar(np.nan, len(index), ---> 451 nan_dtype) 452 arrays.loc[missing] = [v] * missing.sum() 453 ~\Anaconda3\lib\site-packages\pandas\core\dtypes\cast.py in construct_1d_arraylike_from_scalar(value, length, dtype) 1194 else: 1195 if not isinstance(dtype, (np.dtype, type(np.dtype))): -> 1196 dtype = dtype.dtype 1197 1198 # coerce if we have nan for an integer dtype AttributeError: type object 'object' has no attribute 'dtype'
报错原因
这个错误是旧版本distfit库与2.0及以上版本pandas不兼容导致的:旧版distfit在内部创建空DataFrame存储拟合结果时,没有正确处理pandas新版本的dtype校验逻辑,触发了属性不存在的报错。
修复方案
按优先级从高到低选择以下任意一种方法即可:
- 升级distfit到最新版本:这是最稳妥的解决方法,最新版distfit已经修复了该兼容性问题。在终端执行命令
pip install --upgrade distfit,执行完成后重启Jupyter内核,重新运行代码即可正常运行。 - 降级pandas到兼容版本:如果因为项目依赖限制无法升级distfit,可以将pandas降级到1.5.3的稳定兼容版本,执行命令
pip install pandas==1.5.3,完成后重启Jupyter内核即可。 - 临时补丁方案:如果既不能升级distfit也不能降级pandas,可以在导入distfit前加入如下补丁代码,手动修正dtype处理逻辑:
import numpy as np import pandas as pd # 打兼容补丁 from pandas.core import dtypes original_func = dtypes.cast.construct_1d_arraylike_from_scalar def patched_func(value, length, dtype): if dtype is object: dtype = np.dtype('O') return original_func(value, length, dtype) dtypes.cast.construct_1d_arraylike_from_scalar = patched_func # 之后再导入和使用distfit from distfit import distfit
内容的提问来源于stack exchange,提问作者Mark
相关产品推荐
相关产品推荐

