基于铰链损失的SVM分类器梯度下降实现问题排查
多分类SVM实现问题排查
我尝试用Python和Numpy在Jupyter Notebook中从零实现并训练多分类SVM分类器,参考CS231n课程的梯度下降内容,实现了如下SVM类:
class SVM: def __init__(self): self.weights = np.random.randn(len(labels), X_train.shape[1]) * 0.1 self.history = [] def predict(self, X): ''' returns class predictions in np array of size n x num_classes, where n is the number of examples in X ''' #matrix multiplication to apply weights to X bounds = self.weights @ X.T #return the predictions return np.array(bounds).T def loss(self, scores, y, delta=1): '''computes the loss''' #calculate and return the loss for a prediction and corresponding truth label #hinge loss in this case total_loss = 0 #compute loss for each example... for i in range(len(scores)): #extract values for this example scores_of_x = scores[i] label = y[i] correct_score = scores_of_x[label] incorrect_scores = np.concatenate((scores_of_x[:label], scores_of_x[label+1:])) #use the scores for example x to compute the loss at x wj_xi = correct_score #these should be a vector of INCORRECT scores wyi_xi = incorrect_scores #this should be a vector of the CORRECT score wy_xi = wj_xi - wyi_xi + delta #core of the hinge loss formula losses = np.maximum(0, wy_xi) #lower bound the losses at 0 loss = np.sum(losses) #sum the losses #add to the total loss total_loss += loss #return the loss avg_loss = total_loss / len(scores) return avg_loss def gradient(self, scores, X, y, delta=1): '''computes the gradient''' #calculate the loss and the gradient of the loss function #gradient of hinge loss function gradient = np.zeros(self.weights.shape) #calculate the gradient in each example in x for i in range(len(X)): #extract values for this example scores_of_x = scores[i] label = y[i] x = X[i] correct_score = scores_of_x[label] incorrect_scores = np.concatenate((scores_of_x[:label], scores_of_x[label+1:])) # ## ### start by computing the gradient of the weights of the correct classifier ## # wj_xi = correct_score #these should be a vector of INCORRECT scores wyi_xi = incorrect_scores #this should be a vector of the CORRECT score wy_xi = wj_xi - wyi_xi + delta #core of the hinge loss formula losses = np.maximum(0, wy_xi) #lower bound the losses at 0 #get number of nonzero losses, and scale data vector by them to get the loss num_contributing_classifiers = np.count_nonzero(losses) #print(f"Num loss contributors: {num_contributing_classifiers}") g = -1 * x * num_contributing_classifiers #NOTE the -, very important here, doesn't apply to other scores #add the gradient of the correct classifier to the gradient gradient[label] += g #because arrays are 0-indexed, but the labels are 1-indexed # print(f"correct label: {label}") #print(f"gradient:\n{gradient}") # ## ### then, compute the gradient of the weights for each incorrect classifier ## # for j in range(len(scores_of_x)): #skip the correct score, since we already did it if j == label: continue wj_xi = scores_of_x[j] #should be a vector containing the score of the CURRENT classifier wyi_xi = correct_score #should be a vector containing the score of the CORRECT classifier wy_xi = wj_xi - wyi_xi + delta #core of the hinge loss formula loss = np.maximum(0, wy_xi) #lower bound the loss at 0 #get whether this classifier contributed to the loss, and scale the data vector by that to get the gradient contributed_to_loss = 0 if loss > 0: contributed_to_loss = 1 g = x * contributed_to_loss #either times 1 or times 0 #add the gradient of the incorrect classifier to the gradient gradient[j] += g #divide the gradient by number of examples to get the average gradient return gradient / len(X) def fit(self, X, y, epochs = 1000, batch_size = 256, lr=1e-2, verbose=True): #gradient descent loop for epoch in range(epochs): self.history.append({'epoch': epoch}) #create a batch of samples to calculate the gradient #NOTE: this significantly boosts the speed of training indices = np.random.choice(len(X), batch_size, replace=False) X_batch = X.iloc[indices] y_batch = y.iloc[indices] X_batch = X_batch.to_numpy() y_batch = y_batch.to_numpy() #evaluate class scores on training set predictions = self.predict(X_batch) predicted_classes = np.argmax(predictions, axis=1) #compute the loss: average hinge loss loss = self.loss(predictions, y_batch) self.history[-1]['loss'] = loss #compute accuracy on the test set, for an intuitive metric accuracy = np.mean(predicted_classes == y_batch) self.history[-1]['accuracy'] = accuracy #print progress if epoch%50 == 0 and verbose: print(f"Epoch: {epoch} | Loss: {loss} | Accuracy: {accuracy} | LR: {lr} \n") #compute the gradient on the scores assigned by the classifier gradient = self.gradient(predictions, X_batch, y_batch) #backpropagate the gradient to the weights + bias step = gradient * lr #perform a parameter update, in the negative??? direction of the gradient self.weights += step
训练时,损失总体呈下降趋势,但准确率却降至0;同时偶尔会出现损失上升的情况。我知道损失与准确率并非直接相关,但正常情况下准确率应随损失下降而上升,因此怀疑loss和gradient方法存在错误,但无法定位问题点。使用的数据集为鱼类物种采样体重和高度数据。
问题排查与修正方案
1. Hinge Loss公式完全搞反
在loss方法中,变量赋值和核心公式错误:
- 注释标注
wj_xi应为错误类别的分数向量,但实际赋值为正确类别分数correct_score;wyi_xi应为正确类别分数,实际赋值为错误类别分数向量incorrect_scores。 - 核心损失公式
wy_xi = wj_xi - wyi_xi + delta完全颠倒,正确的多分类SVM hinge loss公式应为:对每个错误类别j,计算max(0, s_j - s_y_i + delta),其中s_j是错误类别分数,s_y_i是正确类别分数。
修正后的loss方法片段:
#use the scores for example x to compute the loss at x # 正确赋值:wj_xi是错误类别分数向量,wyi_xi是正确类别分数 wj_xi = incorrect_scores wyi_xi = correct_score # 正确的hinge loss公式:错误分数 - 正确分数 + delta wy_xi = wj_xi - wyi_xi + delta losses = np.maximum(0, wy_xi) loss = np.sum(losses)
2. 梯度计算逻辑对应错误
gradient方法中,正确类别梯度的计算基于错误的loss公式,导致统计的“贡献损失的类别数量”完全错误,进而梯度的符号和数值都偏离正确方向。
修正后的gradient方法中正确类别梯度计算片段:
# 正确的变量赋值与loss公式 wj_xi = incorrect_scores wyi_xi = correct_score wy_xi = wj_xi - wyi_xi + delta losses = np.maximum(0, wy_xi) # 统计满足损失条件的错误类别数量 num_contributing_classifiers = np.count_nonzero(losses) # 正确类别的梯度:-x * 贡献损失的错误类别数量 g = -1 * x * num_contributing_classifiers gradient[label] += g
3. 权重更新方向错误
在fit方法中,梯度下降的更新逻辑错误:原代码执行self.weights += step,其中step = gradient * lr,但梯度是损失函数上升的方向,要最小化损失应该朝着梯度的反方向更新,即减去梯度乘以学习率。
修正后的权重更新代码:
# 正确的梯度下降更新:减去梯度*学习率 self.weights -= lr * gradient
4. 初始化方法依赖外部变量(可选优化)
__init__方法中直接使用外部的labels和X_train变量,导致类的可复用性极差,建议改为传入参数:
def __init__(self, num_classes, num_features): self.weights = np.random.randn(num_classes, num_features) * 0.1 self.history = []
内容的提问来源于stack exchange,提问作者ho88it
相关产品推荐
相关产品推荐

