numdifftools与scipy.optimize.minimize梯度计算结果不一致问题咨询
Let’s break down the possible reasons for this discrepancy, using the details from your output:
1. You’re calculating gradients at different points (most likely cause)
First, double-check that test.refrigeration.params is exactly equal to test.refrigeration.solution.x (the optimal point found by scipy). A common mistake is forgetting to update the params variable with the optimized x value after running minimize.
Looking at your scipy output, the optimal x is:
array([4.14338350e-08, 4.94572523e+00, 8.25733428e+00, 2.05952085e+02])
Print both test.refrigeration.params and test.refrigeration.solution.x side-by-side to confirm they match (including floating-point precision). If they don’t, that’s the root cause—you’re computing gradients at two different locations.
2. Finite difference implementation differences
Scipy and numdifftools use different default settings for numerical gradient calculation, which can lead to divergent results:
- Step size: Scipy’s internal finite difference (used when you don’t provide an analytic Jacobian) uses adaptive step sizes tailored to the problem, while numdifftools uses a fixed default step. For parameters with very small values (like your first parameter, ~4e-8), a poorly chosen step size can introduce massive numerical error.
- Difference type: Scipy may use forward differences for speed, while numdifftools defaults to more accurate (but slower) central differences. Alternatively, if you provided an analytic Jacobian to scipy, its
gradwill be exact (or near-exact), while numdifftools is still approximating via finite differences. - Numerical stability: Scipy’s optimization routines have built-in safeguards for edge cases (like tiny parameter values) that numdifftools might not handle the same way.
To test this, try forcing numdifftools to match scipy’s behavior:
# Use forward differences instead of central nd.Gradient(test.refrigeration.loss_func, method='forward')(test.refrigeration.params) # Adjust step size to match scipy's default eps=1e-8 nd.Gradient(test.refrigeration.loss_func, step=1e-8)(test.refrigeration.params)
3. Confusing target function gradient vs. Lagrangian gradient
Scipy’s output includes two gradient-related fields:
grad: The gradient of your raw loss function at the optimalx.lagrangian_grad: The gradient of the Lagrangian function (loss + constraint penalties) atx, which should be near-zero at convergence (matches your output: ~1e-14 to 1e-6).
Your numdifftools call computes the raw loss function’s gradient, so it should align with scipy’s grad—unless one of the above issues is present. If you accidentally expected to compare against lagrangian_grad, that would explain the mismatch, but your question states you’re comparing to scipy’s grad.
4. Side effects in your loss function
Check if your loss_func has any hidden side effects (e.g., modifying global variables, updating internal state, or relying on dynamic external data) that cause it to return different values when called by scipy vs. numdifftools. For example:
- If
loss_funcmodifiestest.refrigeration.paramsduring calculation, each call could alter the input point for subsequent calculations. - If it depends on cached values that change between runs, the function’s output (and thus its gradient) would differ.
To verify this, call test.refrigeration.loss_func twice in a row with the same params and confirm you get identical results.
内容的提问来源于stack exchange,提问作者Hashemi Emad

