You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

将NumPy数组传入TensorFlow神经网络时维度不匹配问题求助

解决泰坦尼克号挑战赛中TensorFlow形状不匹配的错误

我帮你排查了代码里的几个关键问题,正是这些问题导致了“两个形状的第0维必须相等”的错误,下面是问题分析和修复方案:

核心问题拆解

  1. 标签维度不匹配:tf.nn.softmax_cross_entropy_with_logits要求标签是one-hot编码的二维数组(比如[[1,0],[0,1]]),但你的生存标签是一维的0/1数组(比如[0,1]),和模型输出的二维预测结果形状不兼容。
  2. Tensor与Placeholder混用:你提前把训练数据转成了Tensor,但Placeholder的作用就是动态接收numpy数组,这种混用不仅没必要,还会导致维度对齐错误。
  3. 未处理字符串特征:你保留了Sex列但没转成数值,直接转float会隐性报错,影响后续数据维度。

分步修复方案

1. 先修复数据集处理(Dataset.py)

首先把字符串类型的性别转成数值,同时简化标签获取逻辑:

import pandas
import numpy as np

dataset = pandas.read_csv('train.csv')
# 将性别字符串转成0/1数值
dataset['Sex'] = dataset['Sex'].map({'male': 0, 'female': 1})
# 移除无关列,保留有效特征
dataset2 = dataset.drop(['PassengerId','Survived','Name','Ticket','Fare','Cabin','Embarked'], axis=1)
# 填充缺失值(这里用0,后续可以优化为均值/中位数)
dataset3 = dataset2.fillna(0)
# 获取生存标签并转成float
survival = np.float32(dataset['Survived'])
# 转成模型可用的float32数组
dataset4 = np.float32(dataset3)

2. 修正神经网络训练逻辑(MainCode.py)

重点调整损失函数、Placeholder使用和标签编码:

import tensorflow as tf
import numpy as np
from dataset import dataset4, survival
from sklearn.model_selection import train_test_split

# 分割训练集和测试集
train_x, test_x, train_y, test_y = train_test_split(dataset4, survival, test_size=0.2)

# 将一维标签转成one-hot编码的二维数组,匹配模型输出形状
with tf.Session() as sess:
    train_y_onehot = sess.run(tf.one_hot(train_y.astype(np.int32), depth=2))
    test_y_onehot = sess.run(tf.one_hot(test_y.astype(np.int32), depth=2))

# 定义Placeholder,明确形状:None表示任意样本数,后面是特征数/类别数
x = tf.placeholder(tf.float32, shape=[None, train_x.shape[1]])
y = tf.placeholder(tf.float32, shape=[None, 2])

# 神经网络模型:注意输入层权重的形状要和特征数匹配
def neural_network_model(data):
    n_nodes_hl1 = 10
    n_nodes_hl2 = 10
    n_classes = 2

    hidden_1_layer = {'weights': tf.Variable(tf.random_normal([train_x.shape[1], n_nodes_hl1])),
                      'biases': tf.Variable(tf.random_normal([n_nodes_hl1]))}
    hidden_2_layer = {'weights': tf.Variable(tf.random_normal([n_nodes_hl1, n_nodes_hl2])),
                      'biases': tf.Variable(tf.random_normal([n_nodes_hl2]))}
    output_layer = {'weights': tf.Variable(tf.random_normal([n_nodes_hl2, n_classes])),
                    'biases': tf.Variable(tf.random_normal([n_classes]))}

    l1 = tf.add(tf.matmul(data, hidden_1_layer['weights']), hidden_1_layer['biases'])
    l1 = tf.nn.relu(l1)
    l2 = tf.add(tf.matmul(l1, hidden_2_layer['weights']), hidden_2_layer['biases'])
    l2 = tf.nn.relu(l2)
    output = tf.matmul(l2, output_layer['weights']) + output_layer['biases']
    return output

def train_neural_network():
    prediction = neural_network_model(x)
    # 改用Placeholder传入one-hot标签,和预测结果维度匹配
    cost = tf.reduce_mean(tf.nn.softmax_cross_entropy_with_logits(logits=prediction, labels=y))
    optimizer = tf.train.GradientDescentOptimizer(0.001).minimize(cost)
    hm_epochs = 100

    with tf.Session() as sess:
        sess.run(tf.global_variables_initializer())
        for epoch in range(hm_epochs):
            epoch_loss = 0
            # 直接传入numpy数组,不需要提前转Tensor
            _, c = sess.run([optimizer, cost], feed_dict={x: train_x, y: train_y_onehot})
            epoch_loss += c
            print(f'Epoch {epoch+1}/{hm_epochs}, Loss: {epoch_loss:.4f}')
        
        # 计算准确率:对比预测类别和真实类别
        correct = tf.equal(tf.argmax(prediction, 1), tf.argmax(y, 1))
        accuracy = tf.reduce_mean(tf.cast(correct, 'float'))
        print(f'Accuracy: {accuracy.eval({x: test_x, y: test_y_onehot}):.4f}')

# 启动训练
train_neural_network()

3. 移除冗余的TensorFlowNumpy.py

现在我们不再需要手动转换numpy和Tensor,Placeholder会自动处理numpy数组的传入,这个文件可以直接删除。

额外优化建议

  • 缺失值处理:Age列用0填充不太合理,建议用dataset['Age'].fillna(dataset['Age'].median())代替,提升模型效果。
  • 学习率调整:0.001的学习率可能偏慢,可以尝试0.01或0.1,观察损失下降速度。
  • 小批量训练:全量训练容易过拟合,你可以把训练数据分成小批次传入,让训练更稳定。

内容的提问来源于stack exchange,提问作者5Volts

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 06:58:00