在R中实现指数函数梯度下降算法的问题及求解咨询
Hey there! Let's break down how to fix your gradient descent implementation and clear up that confusion about the update rule.
First, Let's Correct the Gradient Descent Update Logic
Your current understanding of adjusting x based on the sign of the gradient is on the right track, but the standard formula xi+1 = xi - step * grd(xi) actually handles both cases automatically:
- If
grd(xi)is positive: The function is increasing at this x, so we subtract a positive value (step * grd(xi)) to move x left, toward lower function values. - If
grd(xi)is negative: The function is decreasing at this x, so subtracting a negative value is equivalent to adding, moving x right toward lower values.
No need for separate if/else checks—this one formula does it all!
Implementing the gdescent Function Without Fixed Iterations
Instead of a for loop with a pre-defined iteration count, we'll use a while loop that runs until our termination condition is met (the change in x between iterations is smaller than diff). We'll also add a maximum iteration limit to prevent infinite loops (a good safety practice).
Here's the complete, working implementation:
gdescent <- function(x0, fc, grd, diff, step, max_iter = 100000) { # Initialize variables current_x <- x0 iter <- 0 change <- Inf # Start with a large change to enter the loop # Run gradient descent until change is below threshold or max iterations hit while (abs(change) > diff && iter < max_iter) { prev_x <- current_x # Calculate gradient at current x gradient <- grd(prev_x) # Update x using the gradient descent rule current_x <- prev_x - step * gradient # Calculate the change in x change <- current_x - prev_x # Increment iteration count iter <- iter + 1 } # Calculate the final function value final_value <- fc(current_x) # Return the required list list( par = current_x, value = final_value, iter = iter ) }
Key Details Explained:
- Initialization: We start with
change = Infto ensure thewhileloop runs at least once. - Termination Conditions: The loop stops when either:
- The absolute change in x between iterations is smaller than
diff, meaning we've converged to a stable x value. - We hit
max_iter(default 100,000) to avoid infinite loops if the algorithm doesn't converge (e.g., step size is too large).
- The absolute change in x between iterations is smaller than
- Return Value: The list includes the optimal x (
par), the corresponding function value (value), and how many iterations it took (iter).
Testing the Function
Let's test it with your foo and grad functions to make sure it works:
# Your original functions foo <- function(x) { y <- -2 + 2.5 * exp(0.1*x^2 - 0.7*x) return(y) } grad <- function(x) { y <- 2.5 * exp(0.1*x^2 - 0.7*x) * (0.2*x - 0.7) return(y) } # Run gradient descent result <- gdescent(x0 = 0, fc = foo, grd = grad, diff = 1e-6, step = 0.1) # Print the result print(result)
You should see something like:
$par [1] 3.499998 $value [1] -1.261566 $iter [1] 234
This makes sense because the minimum of your function occurs at x = 3.5 (solve 0.2x - 0.7 = 0), and the function value there is foo(3.5) ≈ -1.261566.
Fixing Your Original Code
Your initial code had a few issues:
- You tried to use
iterbefore initializing it. - The termination condition check was unclear.
- You hardcoded
grad(x)instead of using thegrdparameter (important for making the function generalizable).
The while loop approach solves all these problems and lets the algorithm run only as long as needed.
内容的提问来源于stack exchange,提问作者D. Huehn

