You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

从字典创建DataFrame遇ValueError:列数组需为一维的解决方法

RFECV性能曲线DataFrame创建报错及绘图问题

原代码与报错

执行代码

model = ExtraTreesRegressor()      
feature_selector = RFECV(estimator=model, step=1, cv=5, scoring='r2') 
feature_selector.fit(X_train, np.ravel(y_train))
feature_names = X_train.columns
selected_features = feature_names[feature_selector.support_].tolist()
performance_curve = {"Number of Features": list(range(1, len(feature_names) + 1)),
                     "r2": (feature_selector.grid_scores_)}
performance_curve = pd.DataFrame(performance_curve)

报错信息

performance_curve = pd.DataFrame(performance_curve)
Traceback (most recent call last):
  File "C:\Users\user\AppData\Local\Temp\ipykernel_3436\1638829063.py", line 1, in <module>
    performance_curve = pd.DataFrame(performance_curve)
  File "C:\Users\user\anaconda3\lib\site-packages\pandas\core\frame.py", line 636, in __init__
    mgr = dict_to_mgr(data, index, columns, dtype=dtype, copy=copy, typ=manager)
  File "C:\Users\user\anaconda3\lib\site-packages\pandas\core\internals\construction.py", line 502, in dict_to_mgr
    return arrays_to_mgr(arrays, columns, index, dtype=dtype, typ=typ, consolidate=copy)
  File "C:\Users\user\anaconda3\lib\site-packages\pandas\core\internals\construction.py", line 120, in arrays_to_mgr
    index = _extract_index(arrays)
  File "C:\Users\user\anaconda3\lib\site-packages\pandas\core\internals\construction.py", line 661, in _extract_index
    raise ValueError("Per-column arrays must each be 1-dimensional")
ValueError: Per-column arrays must each be 1-dimensional

当前数据结构

字典中Number of Features是长度9的一维列表,r2是形状(9,5)的二维数组(对应9个特征数,每个特征数下5折交叉验证的分数):

{'Number of Features': [1, 2, 3, 4, 5, 6, 7, 8, 9],
 'r2': array([[0.897 , 0.8891, 0.9031, 0.8967, 0.8833],
        [0.889 , 0.8822, 0.8906, 0.8828, 0.8801],
        [0.9468, 0.9388, 0.9411, 0.9448, 0.9401],
        [0.9623, 0.9567, 0.9564, 0.9539, 0.9576],
        [0.9674, 0.962 , 0.9612, 0.9643, 0.9634],
        [0.9958, 0.9939, 0.9925, 0.9944, 0.9928],
        [0.9959, 0.9939, 0.9924, 0.9945, 0.993 ],
        [0.9961, 0.9941, 0.9926, 0.9949, 0.9929],
        [0.9963, 0.9943, 0.9926, 0.995 , 0.993 ]])}

问题原因

scikit-learn版本更新后,RFECV.grid_scores_返回二维数组(每个特征数对应所有交叉验证折的分数),而Pandas要求DataFrame的每列必须是一维数组,直接转换会触发维度不匹配报错。旧版本中grid_scores_仅返回各特征数下的平均分数(一维数组),因此当时可正常运行。

解决方法

方法1:使用平均分数生成简洁曲线

如果只需展示每个特征数对应的平均R²分数,直接计算二维数组的行均值,转为一维数组后创建DataFrame:

# 计算每个特征数的5折平均R²
performance_curve = {
    "Number of Features": list(range(1, len(feature_names) + 1)),
    "r2": feature_selector.grid_scores_.mean(axis=1)  # axis=1计算每行均值,转为一维数组
}
performance_curve = pd.DataFrame(performance_curve)

# 原绘图代码可直接正常运行
sns.lineplot(x = "Number of Features", y = "r2", data = performance_curve,
             color = line_color, lw = 4, ax = ax)
sns.regplot(x = performance_curve["Number of Features"], y = performance_curve["r2"],
            color = marker_colors, fit_reg = False, scatter_kws = {"s": 200}, ax = ax)

方法2:展开所有交叉验证分数(展示全部折的结果)

如果需要展示每个特征数下所有5折的分数分布,可将二维数据展开为一维,同时重复对应特征数:

import pandas as pd
import numpy as np

n_features_list = list(range(1, len(feature_names) + 1))
r2_scores = feature_selector.grid_scores_

# 展开数据:每个折的分数对应一行
expanded_rows = []
for n, scores in zip(n_features_list, r2_scores):
    for score in scores:
        expanded_rows.append({
            "Number of Features": n,
            "r2": score
        })

performance_curve = pd.DataFrame(expanded_rows)

# 绘图:展示所有散点+均值折线
import seaborn as sns
import matplotlib.pyplot as plt

fig, ax = plt.subplots()
# 绘制所有交叉验证的散点
sns.regplot(x="Number of Features", y="r2", data=performance_curve,
            color="gray", fit_reg=False, scatter_kws={"s": 30}, ax=ax)
# 计算并绘制均值折线
mean_r2 = performance_curve.groupby("Number of Features")["r2"].mean().reset_index()
sns.lineplot(x="Number of Features", y="r2", data=mean_r2,
             color="red", lw=3, ax=ax)
plt.show()

推荐方案:使用cv_results_替代弃用的grid_scores_

scikit-learn已标记grid_scores_为弃用,建议使用cv_results_属性直接获取平均分数,更符合新版本规范:

performance_curve = {
    "Number of Features": list(range(1, len(feature_names) + 1)),
    "r2": feature_selector.cv_results_['mean_test_score']  # 直接获取平均测试分数
}
performance_curve = pd.DataFrame(performance_curve)

内容的提问来源于stack exchange,提问作者Progalu

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.19 18:54:55