基于论文1704.04289的超参数优化(公式35)Python实现求助
Hey there! Let's walk through how to implement Equation 35 from Section 7.3 of that paper into your existing logistic regression code. I know you already grasp the theory, so let's jump straight to mapping those paper terms to working Python code.
First, let's recap the key terms from Equation 35 to align with your setup:
Equation 35 (adapted for your logistic regression case):
$$
\mathcal{L}(\theta) = -\log p(y|X,\theta) + \lambda \sum_{k=1}^K \frac{1}{|D_k|} |\theta(D_k)|2^2 + \lambda |\theta{dyn}|_2^2
$$
- $-\log p(y|X,\theta)$: The negative of your existing
log_likelihoodoutput (your current function returns the positive log-likelihood; we flip it to get a loss we can minimize)- $\lambda$: The regularization hyperparameter you'll tune later
- $D_k$: Groups of static parameters (e.g., split your feature weights into logical groups)
- $|D_k|$: Number of parameters in group $D_k$
- $\theta(D_k)$: The subset of model weights belonging to group $D_k$
- $\theta_{dyn}$: Dynamic parameters (typically the bias term, since it's not tied to a specific feature group)
Step 1: Update the Loss Function with Regularization
First, let's rewrite your log-likelihood function to return the regularized negative log-likelihood (the loss $\mathcal{L}$ we need to minimize):
def regularized_nll(features, target, weights, lambda_reg, param_groups, dyn_param_idx): # Calculate base negative log-likelihood (flip the sign of your original log-likelihood) scores = np.dot(features, weights) base_nll = -np.sum(target * scores - np.log(1 + np.exp(scores))) # Calculate static group regularization: λ * sum( (1/|D_k|) * ||θ(D_k)||² ) reg_static = 0.0 for group in param_groups: group_weights = weights[group] group_size = len(group) reg_static += (1 / group_size) * np.sum(group_weights ** 2) reg_static *= lambda_reg # Calculate dynamic parameter regularization: λ * ||θ_dyn||² dyn_weights = weights[dyn_param_idx] reg_dyn = lambda_reg * np.sum(dyn_weights ** 2) # Total loss is the sum of all terms total_loss = base_nll + reg_static + reg_dyn return total_loss
Parameter Breakdown:
lambda_reg: The $\lambda$ from the paper (start with 0.1 if you're testing)param_groups: A list of index lists. For your simulated data,Xbhas columns[bias, x1, x2], so you could use[[1], [2]](each feature weight is its own group) or[[1,2]](both features in one group)dyn_param_idx: The index of your dynamic parameter (almost always the bias term, index 0 in your setup)
Step 2: Update Gradient Descent to Include Regularization
Next, we need to add the regularization terms to the gradient update. The gradient of the regularization terms is straightforward:
- For each static group $D_k$: gradient = $\lambda * (2 / |D_k|) * \theta(D_k)$
- For the dynamic parameter: gradient = $\lambda * 2 * \theta_{dyn}$
Here's the modified training function:
def regularized_log_reg(features, target, num_steps, learning_rate, lambda_reg, param_groups, dyn_param_idx): weights = np.zeros(features.shape[1]) for step in range(num_steps): scores = np.dot(features, weights) predictions = sigmoid(scores) # Base gradient from standard logistic regression (maximizes log-likelihood) error = target - predictions gradient = np.dot(features.T, error) # Calculate regularization gradient (since we're minimizing loss, we subtract this from the base gradient) reg_gradient = np.zeros_like(weights) # Add static group regularization gradients for group in param_groups: group_weights = weights[group] group_size = len(group) reg_gradient[group] += (2 * lambda_reg / group_size) * group_weights # Add dynamic parameter regularization gradient reg_gradient[dyn_param_idx] += 2 * lambda_reg * weights[dyn_param_idx] # Update weights: adjust base gradient by regularization term gradient -= reg_gradient weights += learning_rate * gradient # Print progress every 10k steps if step % 10000 == 0: current_loss = regularized_nll(features, target, weights, lambda_reg, param_groups, dyn_param_idx) print(f"Step {step}, Total Loss: {current_loss:.4f}") return weights
Step 3: Test with Your Simulated Data
Let's run the regularized model with your existing simulated data. We'll use:
- $\lambda = 0.1$ (you can tune this later with cross-validation)
- Separate groups for each feature weight
- Bias term as the dynamic parameter
# Regularization settings lambda_reg = 0.1 param_groups = [[1], [2]] # Feature 1 (index 1) = group 1, Feature 2 (index 2) = group 2 dyn_param_idx = [0] # Bias term is the dynamic parameter # Train the regularized model regularized_weights = regularized_log_reg( Xb, simulated_labels, num_steps=50000, learning_rate=5e-5, lambda_reg=lambda_reg, param_groups=param_groups, dyn_param_idx=dyn_param_idx ) # Compare with unregularized weights (from your original function) original_weights = log_reg(Xb, simulated_labels, num_steps=50000, learning_rate=5e-5) print("\nOriginal Unregularized Weights:", original_weights) print("Regularized Weights:", regularized_weights)
Quick Tips for Alignment with the Paper
- If your parameter grouping is different (e.g., multiple features per group), just adjust the
param_groupslist to include the correct weight indices. - If you don't have dynamic parameters, simply remove the
reg_dynterm from the loss and the corresponding gradient line. - Tune $\lambda$ using cross-validation (split your data into train/validation sets and pick the value that minimizes validation loss).
内容的提问来源于stack exchange,提问作者mkpisk

