You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PyTorch交叉熵损失dtype不匹配问题求助:期望Long却得Float

解决PyTorch系外行星分类模型的CrossEntropyLoss dtype不匹配问题

错误原因

  1. 标签类型错误:PyTorch的CrossEntropyLoss要求目标标签必须是torch.long类型(整数张量),但原代码中把y_train/y_test转成了float类型。
  2. 模型输出设置错误:
    • 手动添加了Softmax层,而CrossEntropyLoss内部已集成LogSoftmax计算,重复使用会导致损失计算异常。
    • 模型输出维度设置为输入特征数(X_train.shape[1]),不符合二分类任务的输出要求(应设为2)。
  3. 设备不匹配:模型移到了GPU/CPU,但输入数据仍在原设备,会导致计算错误。

修复步骤

  1. 修正标签数据类型与取值:
    • 开普勒数据集的LABEL取值为1(无行星)和2(有行星),需先减1转为0/1的标准分类标签。
    • 将标签张量转换为torch.long类型。
  2. 调整模型结构:
    • 移除Softmax层。
    • 将输出层维度改为2(对应二分类任务)。
  3. 统一设备:将输入数据移至与模型相同的设备。

完整修正代码

import pandas as pd
import torch as T
import torch.nn as nn
import torch.optim as opt
from sklearn.model_selection import train_test_split
from sklearn.preprocessing import MinMaxScaler

# 加载数据
train_df = pd.read_csv("../csvs/Space Travel/Exoplanet Hunting in Deep Space/1/exoTrain.csv")
test_df = pd.read_csv("../csvs/Space Travel/Exoplanet Hunting in Deep Space/1/exoTest.csv")

# 预处理特征与标签
X_train_df = train_df.drop(["LABEL"], axis=1).values
y_train_df = train_df["LABEL"].values.reshape(-1,1).squeeze()
# 修正标签:1→0,2→1
y_train_df = y_train_df - 1

# 划分训练测试集
X_train, X_test, y_train, y_test = train_test_split(X_train_df,
                                                   y_train_df,
                                                   test_size=0.2,
                                                   train_size=0.8,
                                                   shuffle=True,
                                                   random_state=123)

# 特征归一化
sc = MinMaxScaler()
X_train = sc.fit_transform(X_train)
X_test = sc.transform(X_test)

# 转换为PyTorch张量,修正数据类型
X_train = T.from_numpy(X_train).float()
X_test = T.from_numpy(X_test).float()
# 标签转为long类型
y_train = T.from_numpy(y_train).long()
y_test = T.from_numpy(y_test).long()

# 定义模型
class Exoplanet_AI(nn.Module):
    def __init__(self, input_dims=X_train.shape[1], hidden_units=125, output_dims=2):
        super().__init__()
        self.activation = nn.LeakyReLU()
        self.droprate = 0.2
        self.dropout = nn.Dropout(p=self.droprate)
        self.flaten = nn.Flatten()
        
        self.ll1 = nn.Linear(in_features=input_dims, out_features=hidden_units)
        self.ll2 = nn.Linear(in_features=hidden_units, out_features=hidden_units)
        self.ll3 = nn.Linear(in_features=hidden_units, out_features=hidden_units)
        # 输出层维度改为2(二分类)
        self.ll4 = nn.Linear(in_features=hidden_units, out_features=output_dims)
        
    def forward(self, X):
        X = self.flaten(X)
        X = self.activation(self.ll1(X))
        X = self.activation(self.ll2(X))
        X = self.dropout(X)
        X = self.activation(self.ll3(X))
        X = self.ll4(X)  # 移除Softmax,CrossEntropyLoss内部已处理
        return X

# 训练测试类
class train_and_testing():
    def __init__(self):
        self.device = T.device("cuda:0" if T.cuda.is_available() else "cpu")
        self.lr = 1e-3
        self.epochs = 20
        
        # 将数据移至目标设备
        self.X_train = X_train.to(self.device)
        self.X_test = X_test.to(self.device)
        self.y_train = y_train.to(self.device)
        self.y_test = y_test.to(self.device)
        
        self.model = Exoplanet_AI().to(self.device)
        self.criterion = opt.Adam(params=self.model.parameters(), lr=self.lr)
        self.loss_fn = nn.CrossEntropyLoss()
        
    def parameters(self):
        return self.model.state_dict()
        
    def train_loop(self):
        current_loss = 0.0
        for i in range(self.epochs):
            self.model.train()
            
            forward_pass = self.model(self.X_train)
            # 计算损失
            loss = self.loss_fn(forward_pass, self.y_train)
            
            loss.backward()
            self.criterion.step()
            self.criterion.zero_grad()
            
            current_loss += loss.item()
            print(f"Epoch {i} | Loss: {loss.item():.4f}")

# 执行训练
trainer = train_and_testing()
trainer.train_loop()

关键说明

  • 标签处理:必须将标签转为long类型,且类别索引从0开始,否则CrossEntropyLoss会报错。
  • Softmax层:CrossEntropyLoss = LogSoftmax + NLLLoss,手动添加Softmax会导致损失计算逻辑错误。
  • 设备统一:模型和输入数据必须在同一设备(CPU/GPU)上,否则会出现张量设备不匹配的错误。

内容的提问来源于stack exchange,提问作者user24470825

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.25 20:12:18