TensorFlow中ScipyOptimizerInterface的Args参数含义咨询
Understanding
ScipyOptimizerInterface L-BFGS-B Parameters Great question! Since tf.contrib.opt.ScipyOptimizerInterface acts as a wrapper for Scipy's L-BFGS-B optimizer, all the keys in the options dictionary directly correspond to the parameters of Scipy's scipy.optimize.minimize method when using the L-BFGS-B solver. Here's a clear breakdown of each parameter you listed:
maxiter: The maximum number of full iterations the L-BFGS-B algorithm will run. An "iteration" here refers to a complete loop of the optimization process (not just a single function evaluation). Your setting of 50000 means the optimizer will halt after 50,000 iterations even if convergence isn't fully met.maxfun: The maximum number of times your cost function can be evaluated. This is a guardrail to prevent excessive computation—if the optimizer hits this limit before reachingmaxiteror convergence criteria, it will stop. Your 50000 value ensures the function isn't called more than 50k times.ftol: The relative tolerance for changes in the cost function value. The optimizer stops when the relative difference between consecutive cost function values drops below this threshold (calculated as|f(k+1) - f(k)| <= ftol * (|f(k)| + ftol)). A value of 1e-15 sets a very strict bar for convergence based on function value stability.gtol: The absolute tolerance for the gradient magnitude. The optimizer stops when the norm of the gradient vector falls below this threshold—this signals we're near a local minimum where the gradient is close to zero. Your 1e-11 setting means the gradient has to be extremely small before stopping.eps: The perturbation size used for numerical gradient calculations (only relevant if you aren't providing an analytical gradient, though TensorFlow usually handles analytical gradients automatically). A value of 1e-15 means using a tiny perturbation to estimate gradients when needed.maxls: The maximum number of steps allowed during the line search phase of each iteration. L-BFGS-B uses line search to find the optimal step size for each update; this parameter caps how many step size attempts the optimizer will make. Your 100 setting gives the line search plenty of room to find a good step size.
For context, your parameter choices are set to very strict convergence criteria, which makes sense if you need high precision in your optimization results.
内容的提问来源于stack exchange,提问作者 George Smith
相关产品推荐
相关产品推荐

