You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

构建DataFrame与相关矩阵报错:数组维度及空矩阵问题求助

问题描述

我正在进行一个数据项目,作为该领域新手,遇到了变量类型相关问题,希望得到帮助。我有8个归一化数组,想要将它们存入DataFrame以构建相关矩阵,但出现错误:

ValueError: Per-column arrays must each be 1-dimensional

我尝试重塑数组但无效,于是检查数组形状:

print(date.shape,normalised_snp.shape,normalised_twybp.shape,normalised_USInflation.shape,normalised_USGDP.shape,normalised_USInterest.shape,normalised_GlobalInflation.shape,normalised_GlobalGDP.shape)

输出结果为:

(4220, 1) (4220, 1) (4220, 1) (4220, 1) (4220, 1) (4220, 1) (4220, 1) (4220, 1)

之后我将数组转为列表并构建DataFrame:

normalised_snp = normalised_snp.tolist()
normalised_tybp = normalised_tybp.tolist()
normalised_twybp = normalised_twybp.tolist()
normalised_USInflation = normalised_USInflation.tolist()
normalised_USGDP = normalised_USGDP.tolist()
normalised_USInterest = normalised_USInterest.tolist()
normalised_GlobalInflation = normalised_GlobalInflation.tolist()
normalised_GlobalGDP = normalised_GlobalGDP.tolist()
alldata = pd.DataFrame({'S&P 500 Price':normalised_snp,
                        '10 Year Bond Price': normalised_tybp,
                        '2 Year Bond Price' : normalised_twybp,
                        'US Inflation' : normalised_USInflation,
                        'US GDP' : normalised_USGDP,
                        'US Insterest' : normalised_USInterest,
                        'Global Inflation Rate' : normalised_GlobalInflation,
                        'Global GDP' : normalised_GlobalGDP})

随后构建相关矩阵:

correlation_matrix = alldata.corr()
print(correlation_matrix)

此时无报错,但相关矩阵为空:

Empty DataFrame
Columns: []
Index: []

请问问题是否由列表类型导致?若如此,如何解决用数组构建DataFrame时出现的ValueError?

解决方案

问题根源

  1. 最初的ValueError是因为数组是2维形状(4220,1),而pandas要求DataFrame的每一列必须是1维数组,直接传入2维数组会被识别为包含子数组的列表,不符合列数据要求。
  2. 转成列表后出现空DataFrame,是因为2维数组转列表后变成[[x1], [x2], ..., [xn]]的嵌套结构,pandas无法正确解析为有效列数据,导致构建的DataFrame无效。

具体解决方法

方法一:将数组转为1维(推荐)

用numpy的ravel()或flatten()方法把2维数组转成1维,再构建DataFrame:

import pandas as pd
import numpy as np

# 将所有2维数组转为1维
normalised_snp = normalised_snp.ravel()
normalised_tybp = normalised_tybp.ravel()
normalised_twybp = normalised_twybp.ravel()
normalised_USInflation = normalised_USInflation.ravel()
normalised_USGDP = normalised_USGDP.ravel()
normalised_USInterest = normalised_USInterest.ravel()
normalised_GlobalInflation = normalised_GlobalInflation.ravel()
normalised_GlobalGDP = normalised_GlobalGDP.ravel()

# 构建DataFrame(修正原代码中的拼写错误:Insterest → Interest)
alldata = pd.DataFrame({
    'S&P 500 Price': normalised_snp,
    '10 Year Bond Price': normalised_tybp,
    '2 Year Bond Price': normalised_twybp,
    'US Inflation': normalised_USInflation,
    'US GDP': normalised_USGDP,
    'US Interest': normalised_USInterest,
    'Global Inflation Rate': normalised_GlobalInflation,
    'Global GDP': normalised_GlobalGDP
})

# 生成并打印相关矩阵
correlation_matrix = alldata.corr()
print(correlation_matrix)

方法二:直接切片取第0列

如果数组是2维结构,也可以通过索引直接提取1维数据:

normalised_snp = normalised_snp[:, 0]
# 其他数组执行同样的切片操作后,再构建DataFrame

方法三:修复嵌套列表(不推荐)

如果坚持用列表,需要把嵌套列表展开为一维列表:

normalised_snp = [x[0] for x in normalised_snp.tolist()]
# 其他数组做同样的列表推导处理

额外提示

原代码中US Insterest存在拼写错误,改为US Interest可避免后续索引或分析时出现问题。

内容的提问来源于stack exchange,提问作者samet.bnc

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.10 07:20:48