求助:解决Cannot describe a DataFrame without columns错误
解决Jupyter Notebook中“Cannot describe a DataFrame without columns”报错的稳定方案
问题背景
两天前完成某数据集的简短数据分析,今日基于该工作开展新项目时复制部分代码,新项目运行正常,但旧项目出现ValueError: Cannot describe a DataFrame without columns错误。此前通过新建Notebook临时解决,现需无需重复新建Notebook的稳定方案。
工作环境:Jupyter Notebook、Python 3.6(虚拟环境)、Linux 22.04
相关代码及报错信息
categorical_features = dtype[dtype == 'object'].index readable_df[numerical_features].describe() # Split features into categorical and numerical, print numerical dtype = readable_df.dtypes numerical_features = dtype[dtype == 'int64'].index categorical_features = dtype[dtype == 'object'].index readable_df[numerical_features].describe()
报错堆栈:
--------------------------------------------------------------------------- ValueError Traceback (most recent call last) Input In [7], in <cell line: 6>() 3 numerical_features = dtype[dtype == 'int64'].index 4 categorical_features = dtype[dtype == 'object'].index ----> 6 readable_df[numerical_features].describe() File ~/python-env/python-env/lib/python3.9/site-packages/pandas/core/generic.py:10227, in NDFrame.describe(self, percentiles, include, exclude, datetime_is_numeric) 9978 @final 9979 def describe( 9980 self: NDFrameT, (...) 9984 datetime_is_numeric=False, 9985 ) -> NDFrameT: 9986 """ 9987 Generate descriptive statistics. 9988 (...) 10225 max NaN 3.0 10226 """ > 10227 return describe_ndframe( 10228 obj=self, 10229 include=include, 10230 exclude=exclude, 10231 datetime_is_numeric=datetime_is_numeric, 10232 percentiles=percentiles, 10233 ) File ~/python-env/python-env/lib/python3.9/site-packages/pandas/core/describe.py:87, in describe_ndframe(obj, include, exclude, datetime_is_numeric, percentiles) 82 describer = SeriesDescriber( 83 obj=cast("Series", obj), 84 datetime_is_numeric=datetime_is_numeric, 85 ) 86 else: --> 87 describer = DataFrameDescriber( 88 obj=cast("DataFrame", obj), 89 include=include, 90 exclude=exclude, 91 datetime_is_numeric=datetime_is_numeric, 92 ) 94 result = describer.describe(percentiles=percentiles) 95 return cast(NDFrameT, result) File ~/python-env/python-env/lib/python3.9/site-packages/pandas/core/describe.py:164, in DataFrameDescriber.__init__(self, obj, include, exclude, datetime_is_numeric) 161 self.exclude = exclude 163 if obj.ndim == 2 and obj.columns.size == 0: --> 164 raise ValueError("Cannot describe a DataFrame without columns") 166 super().__init__(obj, datetime_is_numeric=datetime_is_numeric) ValueError: Cannot describe a DataFrame without columns
报错原因
报错核心是readable_df[numerical_features]返回了无列名的空DataFrame,常见诱因:
- Jupyter内核状态污染:旧Notebook的内核中,
readable_df或numerical_features被后续单元格代码意外覆盖,或之前运行的单元格导致变量状态异常 - 数据加载异常:旧Notebook中
readable_df的加载路径(如相对路径)失效,或文件本身被修改,导致加载后的DataFrame无int64类型列,进而numerical_features为空数组 - 内核环境不匹配:旧Notebook误使用了其他虚拟环境的内核,不同pandas版本对空索引的处理逻辑存在差异
稳定解决方案
1. 重置内核并重新运行全量代码
直接清空当前内核的所有变量状态,从头执行所有代码:
- 在Jupyter界面点击菜单栏
Kernel->Restart & Run All
2. 增加变量校验逻辑
在调用describe()前添加校验,避免空索引导致的报错:
# Split features into categorical and numerical, print numerical dtype = readable_df.dtypes numerical_features = dtype[dtype == 'int64'].index categorical_features = dtype[dtype == 'object'].index # 校验数值特征是否存在 if len(numerical_features) == 0: print("警告:未找到int64类型的数值特征") # 可选:尝试匹配float64类型特征 numerical_features = dtype[dtype == 'float64'].index if len(numerical_features) == 0: print("无可用数值特征,跳过describe操作") else: print(readable_df[numerical_features].describe()) else: print(readable_df[numerical_features].describe())
3. 确认Notebook使用正确的虚拟环境内核
- 检查界面右上角显示的内核名称,确认是你创建的Python3.6虚拟环境
- 若内核不对,点击
Kernel->Change Kernel选择对应环境 - 若虚拟环境未添加到Jupyter内核,激活虚拟环境后执行:
pip install ipykernel python -m ipykernel install --user --name=python3.6-env --display-name="Python 3.6 (虚拟环境)"
4. 固化数据加载路径
将数据加载的相对路径改为绝对路径,避免路径失效问题:
import pandas as pd # 替换为你的数据集绝对路径 readable_df = pd.read_csv("/home/your_username/your_dataset_path/data.csv")
内容的提问来源于stack exchange,提问作者Solaris
相关产品推荐
相关产品推荐

