如何在Pandas DataFrame单元格存储NumPy数组?为何部分代码报错?
多列DataFrame存储NumPy数组报错,单列正常的原因与解决方法
问题复现
以下代码执行失败:
import numpy as np import pandas as pd bnd1 = np.random.rand(74,8) bnd2 = np.random.rand(74,8) df = pd.DataFrame(columns = ["val", "unit"]) df.loc["bnd"] = [bnd1, "N/A"] df.loc["bnd"] = [bnd2, "N/A"]
报错信息:
ValueError: setting an array element with a sequence. The requested array has an inhomogeneous shape after 1 dimensions. The detected shape was (2,) + inhomogeneous part.
完整回溯信息:
> --------------------------------------------------------------------------- AttributeError Traceback (most recent call > last) File > ~/anaconda3/envs/py38mats/lib/python3.8/site-packages/numpy/core/fromnumeric.py:3185, > in ndim(a) 3184 try: > -> 3185 return a.ndim 3186 except AttributeError: > > AttributeError: 'list' object has no attribute 'ndim' > > During handling of the above exception, another exception occurred: > > ValueError Traceback (most recent call > last) Cell In[10], line 8 > 6 df = pd.DataFrame(columns = ["val", "unit"]) > 7 df.loc["bnd"] = [bnd1, "N/A"] > ----> 8 df.loc["bnd"] = [bnd2, "N/A"] > > File > ~/anaconda3/envs/py38mats/lib/python3.8/site-packages/pandas/core/indexing.py:849, > in _LocationIndexer.__setitem__(self, key, value) > 846 self._has_valid_setitem_indexer(key) > 848 iloc = self if self.name == "iloc" else self.obj.iloc > --> 849 iloc._setitem_with_indexer(indexer, value, self.name) > > File > ~/anaconda3/envs/py38mats/lib/python3.8/site-packages/pandas/core/indexing.py:1835, > in _iLocIndexer._setitem_with_indexer(self, indexer, value, name) > 1832 # align and set the values 1833 if take_split_path: 1834 > # We have to operate column-wise > -> 1835 self._setitem_with_indexer_split_path(indexer, value, name) 1836 else: 1837 self._setitem_single_block(indexer, > value, name) > > File > ~/anaconda3/envs/py38mats/lib/python3.8/site-packages/pandas/core/indexing.py:1872, > in _iLocIndexer._setitem_with_indexer_split_path(self, indexer, value, > name) 1869 if isinstance(value, ABCDataFrame): 1870 > self._setitem_with_indexer_frame_value(indexer, value, name) > -> 1872 elif np.ndim(value) == 2: 1873 # TODO: avoid np.ndim call in case it isn't an ndarray, since 1874 # that will > construct an ndarray, which will be wasteful 1875 > self._setitem_with_indexer_2d_value(indexer, value) 1877 elif > len(ilocs) == 1 and lplane_indexer == len(value) and not > is_scalar(pi): 1878 # We are setting multiple rows in a single > column. > > File <__array_function__ internals>:200, in ndim(*args, **kwargs) > > File > ~/anaconda3/envs/py38mats/lib/python3.8/site-packages/numpy/core/fromnumeric.py:3187, > in ndim(a) 3185 return a.ndim 3186 except AttributeError: > -> 3187 return asarray(a).ndim > > ValueError: setting an array element with a sequence. The requested array has an inhomogeneous shape after 1 dimensions. The detected shape was (2,) + inhomogeneous part.
但单列DataFrame的代码可以正常运行:
import numpy as np import pandas as pd bnd1 = np.random.rand(74,8) bnd2 = np.random.rand(74,8) df = pd.DataFrame(columns = ["val"]) df.loc["bnd"] = [bnd1] df.loc["bnd"] = [bnd2]
使用版本:pandas 2.0.3、numpy 1.24.4
原因分析
- 单列场景:给单列赋值
[bnd1]时,Pandas会将整个numpy数组作为单个对象存入单元格,不会尝试将列表转换为多维数组,因此执行正常。 - 多列场景:给多列赋值
[bnd2, "N/A"]时,Pandas内部会尝试将列表转换为numpy数组以匹配列数。但列表包含二维numpy数组和字符串两种不同类型/形状的元素,numpy无法创建形状均匀的数组,从而抛出ValueError。
解决方案
方法1:逐个列单独赋值
避免一次性给整行赋值,分别操作每个列:
import numpy as np import pandas as pd bnd1 = np.random.rand(74,8) bnd2 = np.random.rand(74,8) df = pd.DataFrame(columns = ["val", "unit"]) df.loc["bnd", "val"] = bnd1 df.loc["bnd", "unit"] = "N/A" df.loc["bnd", "val"] = bnd2 df.loc["bnd", "unit"] = "N/A"
方法2:使用pd.Series进行整行赋值
通过pd.Series明确元素类型,避免Pandas自动转换为numpy数组:
import numpy as np import pandas as pd bnd1 = np.random.rand(74,8) bnd2 = np.random.rand(74,8) df = pd.DataFrame(columns = ["val", "unit"]) df.loc["bnd"] = pd.Series([bnd1, "N/A"], index=["val", "unit"]) df.loc["bnd"] = pd.Series([bnd2, "N/A"], index=["val", "unit"])
方法3:初始化时指定dtype=object
让DataFrame以对象类型存储单元格,确保numpy数组可直接存入:
import numpy as np import pandas as pd bnd1 = np.random.rand(74,8) bnd2 = np.random.rand(74,8) df = pd.DataFrame(columns = ["val", "unit"], dtype=object) df.loc["bnd"] = [bnd1, "N/A"] df.loc["bnd"] = [bnd2, "N/A"]
内容的提问来源于stack exchange,提问作者R Walser
相关产品推荐
相关产品推荐

