为何np.mean计算神经网络准确率比Python循环慢且结果不一致?
我在构建一个二分类神经网络时,发现运行速度极慢,排查后定位到用np.mean计算准确率的步骤平均耗时0.55秒,而自己用循环实现的计算仅需0.01秒以内,且两者输出的准确率结果还不一致。
相关代码
原准确率计算代码
predictions = self.neural_network.propagate(X_train) rounded_predictions = (predictions > 0.5).astype(int) accuracy = np.mean(rounded_predictions == y_train)
其中X_train形状为(16000,16),y_train形状为(16000,)。
自定义循环实现的准确率计算
def compute_accuracy_score(y_train, predictions): """ Calculates the accuracy score of the predictions made by the network. If a prediction matches its true label, it increments the count of correct predictions. The accuracy score is the correct predictions divided by the total predictions. """ num_samples = len(y_train) correct_predictions = 0 # If the prediction is correct, increment the count of correct predictions for true_label, predicted_label in zip(y_train, predictions): if true_label == predicted_label: correct_predictions += 1 # Compute accuracy as the ratio of correct predictions to total number of samples accuracy = correct_predictions / num_samples return accuracy
测试运行时间的代码
def calculate_fitness(self): predictions = self.neural_network.propagate(X_train) rounded_predictions = (predictions > 0.5).astype(int) start_time = time.time() accuracy1 = compute_accuracy_score(y_train, rounded_predictions) end_time = time.time() print('Time to calculate fitness using compute accuracy:', end_time - start_time, 'seconds') start_time = time.time() accuracy = np.mean(rounded_predictions == y_train) end_time = time.time() print('Time to calculate fitness using np.mean:', end_time - start_time, 'seconds') print('Accuracy:', accuracy, 'Accuracy1:', accuracy1) self.fitness = round(float(accuracy), 4)
测试结果
Time to calculate fitness using compute accuracy: 0.022114276885986328 seconds Time to calculate fitness using np.mean: 0.6563560962677002 seconds Accuracy: 0.559151625 Accuracy1: 0.6015625 Time to calculate fitness using compute accuracy: 0.021242141723632812 seconds Time to calculate fitness using np.mean: 0.6553714275360107 seconds Accuracy: 0.441705 Accuracy1: 0.398875
差异原因分析
1. 运行速度差异:数组形状不匹配引发的超大临时数组开销
最核心的问题是rounded_predictions与y_train形状不一致,比如rounded_predictions是(16000,1)而y_train是(16000,)。此时执行rounded_predictions == y_train时,numpy会触发广播机制,将两个数组强制扩展为(16000,16000)的超大布尔数组(包含2.56亿个元素),创建这个数组以及计算全局均值的内存和计算开销极大,直接导致耗时飙升。
而自定义循环通过zip逐个遍历元素,完全避免了超大临时数组的创建和内存拷贝,因此速度快得多。即使形状看似匹配,若数组内存布局不连续(比如是切片结果)或数据类型不一致,numpy的向量操作也会产生额外转换开销,而循环遍历不受这些因素影响。
2. 结果不一致:广播后的均值计算逻辑完全错误
当rounded_predictions与y_train形状不匹配时,rounded_predictions == y_train生成的是广播后的超大数组,此时np.mean计算的是所有元素的全局均值,而非每个样本预测与真实标签的匹配率,完全偏离了准确率的计算逻辑。
而自定义循环中,zip(y_train, predictions)会按顺序逐个匹配样本的真实标签和预测值(即使预测值是长度为1的数组,Python在if判断中会自动将单元素布尔数组转为布尔值),因此计算的是正确的样本匹配率。
快速解决方案
先统一数组形状,确保rounded_predictions与y_train维度一致,再计算准确率:
# 方案1:展平预测数组后计算均值 accuracy = np.mean(rounded_predictions.flatten() == y_train) # 方案2:用求和替代均值,效率更高 accuracy = np.sum(rounded_predictions.flatten() == y_train) / len(y_train)
内容的提问来源于stack exchange,提问作者Yair

