TensorFlow整数被3整除分类模型失效问题求助
整数能否被3整除分类模型失效问题分析
背景
学习TensorFlow时,有一段可对整数能否被2整除进行分类的代码,仅添加numpy导入语句后即可正常运行,准确率达100%。仅修改一行代码将任务改为判断整数能否被3整除后,模型随即失效:始终预测类别1,准确率仅为0.67。尝试更换损失函数、添加隐藏层、更换激活函数等方法均无效,需明确问题原因。
代码修改对比
原代码中生成标签的行:
Y.append( to_categorical(v%2, 2) )
修改后的代码行:
Y.append( to_categorical(0 if v%3 == 0 else 1, 2) )
完整代码
import numpy as np from keras.models import Sequential from keras.layers import Dense from keras.utils import to_categorical # Helper function to convert a number # to its fixed width binary representation def conv(x): a = format(x, '032b') l = list(str(a)) l = np.array(list(map(int, l))) return l # input data data = [conv(i) for i in range(100000)] X = np.array(data) Y= list() # empty list of results for v in range(100000): Y.append( to_categorical(0 if v%3 == 0 else 1, 2) ) Y = np.array(Y) # we need np.array # Sequential is a fully connected network model = Sequential() # 32 inputs and 1 neuron in the first layer (hidden layer) model.add(Dense(1, input_dim=32, activation='relu')) # 2 output layer model.add(Dense(2, activation='sigmoid')) model.compile(loss='binary_crossentropy', optimizer='adam', metrics=['accuracy']) # epochs is the number of times to retrain over the same data set # batch_size is how may elements to process in parallel at one go model.fit(X, Y, epochs=5, batch_size=100, verbose=1) weights, biases = model.layers[0].get_weights() print("weights",weights.size, weights, "biases", biases) model.summary()
问题原因
- 数据分布不平衡:0到99999的整数中,仅约1/3的数能被3整除(类别0),剩下2/3都属于类别1。模型在训练时会倾向于预测占比更高的类别,以此降低整体损失,这直接导致准确率停留在多数类的占比0.67左右。
- 任务复杂度与模型容量不匹配:判断能否被2整除仅需看二进制的最后一位,属于线性可分问题,单个神经元就能轻松学到规律;但判断能否被3整除需要计算二进制各位的加权和(权值按2^i mod3循环,即1、2、1、2...),这是一个非线性的复杂规则。原模型仅用1个隐藏神经元的ReLU层,容量严重不足,无法学习到这种复杂模式。
- 损失函数与输出配置不匹配:当前输出层用2个神经元的sigmoid激活,却搭配
binary_crossentropy损失函数。binary_crossentropy适用于单输出的二分类任务,而多输出的分类任务应搭配categorical_crossentropy和softmax激活,这种配置错误也会干扰模型的学习过程。
解决建议
- 平衡数据分布:在
model.fit()中添加class_weight={0:2, 1:1}参数,给少数类更高的损失权重;或通过过采样少数类、欠采样多数类调整数据分布。 - 提升模型容量:增加隐藏层神经元数量(比如将第一个Dense层改为32个神经元),或添加额外的隐藏层,让模型具备学习复杂规则的能力。
- 修正损失与激活配置:将输出层激活改为
softmax,损失函数换成categorical_crossentropy;或者将标签改为一维的0/1值(无需to_categorical),输出层用单个神经元搭配sigmoid激活,损失函数保持binary_crossentropy。
内容的提问来源于stack exchange,提问作者zomega
相关产品推荐
相关产品推荐

