Keras中model.evaluate与model.predict计算指标结果不一致原因咨询
Hey there! Let's dig into why you're seeing that small gap between the DICE score from model.evaluate() and your manual post-prediction calculation. Your hunch about tf.float64 conversion is on the right track, but there are a couple more factors at play here:
1. Type Conversion & Precision Differences
Most Keras/TensorFlow models default to using float32 for computations—it's faster and uses less memory. When you run model.evaluate(), it calculates the DICE score using this default precision.
In your manual code, you're casting each per-sample DICE result to tf.float64 before summing. Even though this is higher precision, converting between float32 and float64 during intermediate steps (like summing intersections) can introduce tiny rounding differences. Over 24 samples, these small shifts add up just enough to create the discrepancy you're seeing.
2. Calculation Approach: Per-Sample Average vs. Global Statistic
model.evaluate() computes the DICE score for the entire batch (your full 24-sample dataset, since batch_size=24) in one go: it calculates the total intersection across all samples, total sum of true masks, total sum of predicted masks, then applies the DICE formula once.
Your manual code, however, calculates the DICE score for each individual sample, converts each to float64, sums those per-sample scores, then averages them. Mathematically, these two approaches should be equivalent—but in practice, floating-point arithmetic has cumulative precision errors. Summing 24 individual DICE values (each with their own rounding) vs. computing from global totals will yield slightly different results.
3. Session Management Issues
You're creating a new tf.Session() for every sample in your loop, and you're missing parentheses on tf.Session().close (it should be tf.Session().close()). While this isn't the primary cause of the discrepancy, repeatedly creating and destroying sessions can introduce subtle inconsistencies in computation contexts. It's better to use a single session for all calculations.
How to Verify & Fix
To confirm these causes, try adjusting your manual calculation to match model.evaluate()'s logic more closely:
# Accumulate global statistics instead of per-sample DICE scores total_intersection = 0.0 total_true_sum = 0.0 total_pred_sum = 0.0 # Use a single session for all computations with tf.Session() as sess: for y_i in range(len(y)): # Flatten masks once per sample y_true_flat = K.flatten(y[y_i]) y_pred_flat = K.flatten(preds[y_i]) # Compute values in the session intersection = sess.run(K.sum(y_true_flat * y_pred_flat)) true_sum = sess.run(K.sum(y_true_flat)) pred_sum = sess.run(K.sum(y_pred_flat)) # Accumulate totals total_intersection += intersection total_true_sum += true_sum total_pred_sum += pred_sum # Calculate overall DICE score (matches model.evaluate()'s approach) dice_result = (2 * total_intersection) / (total_true_sum + total_pred_sum) print(dice_result)
If you run this, you'll likely see the score align much closer to model.evaluate()'s output. Additionally, try removing the tf.float64 cast—using the default float32 will further reduce the precision gap.
Final Takeaway
The tiny difference you're seeing is almost certainly a combination of:
- Floating-point precision errors from per-sample averaging vs. global statistic calculation
- Type conversion between
float32(used by the model/evaluate) andfloat64(used in your manual code)
These are normal in numerical computing, and the magnitude of the difference (~0.00015) is negligible for most semantic segmentation tasks.
内容的提问来源于stack exchange,提问作者Thb

