为何在Python函数内直接修改numpy矩阵theta无效?
theta elements directly work inside my function? Let's break down what's happening here and why your direct assignment to theta isn't behaving as expected inside the function, even though it works outside.
First, the core issue in your modified code
When you removed theta = temp and started updating theta[0,0]/theta[0,1] directly, you also commented out theta_trans = theta.T inside the loop. While theta_trans is technically a view of theta's transpose (so it should update automatically when theta changes), there's a bigger logic flaw breaking your code:
In your loop, you calculate theta[0,1] using the already-updated theta[0,0] from the same iteration. Gradient descent requires using the original state of theta (from the start of the loop) to compute all parameter updates—updating one parameter mid-loop and using that new value for the next calculation skews the gradient and prevents meaningful changes to theta.
Why it works outside the function
When you run:
theta = np.matrix(np.array([0,0])) theta[0,0] = theta[0,0] - 1 print(theta)
You're modifying a single element without any dependent calculations relying on the old theta values afterward. There's no loop or cross-parameter dependencies to mess with the update, so it works exactly as you expect.
Fixing the function
To make direct updates to theta work correctly (without using the temp matrix), you need to calculate both parameter updates using the original theta values from the start of each iteration, then apply both changes at once:
def gradient_descent(X, y, theta, iterations, alpha): m = len(X) for j in range(iterations): # Capture the current theta state to use for both updates current_theta0 = theta[0,0] current_theta1 = theta[0,1] # Calculate hypothesis using the unmodified theta hyp = np.dot(X, theta.T) - y # Compute both updates using the original theta values update0 = current_theta0 - ((alpha / m) * np.sum(np.multiply(hyp, X[:,0]))) update1 = current_theta1 - ((alpha / m) * np.sum(np.multiply(hyp, X[:,1]))) # Apply both updates simultaneously theta[0,0] = update0 theta[0,1] = update1 return theta
Key takeaways
- Gradient descent depends on using the full, unmodified state of your parameters at the start of each iteration to compute accurate updates. Never update one parameter and reuse it for another update in the same loop.
- Mutable objects like numpy matrices are modified in-place inside functions—your direct assignments weren't "invalid" on their own, but your loop logic prevented them from producing meaningful changes.
- Explicitly capturing parameter values at the start of each iteration avoids subtle bugs from partial updates.
内容的提问来源于stack exchange,提问作者BLP

