You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于论文1704.04289的超参数优化(公式35)Python实现求助

Implementing Regularized Negative Log-Likelihood from arXiv:1704.04289 (Section 7.3, Equation 35)

Hey there! Let's walk through how to implement Equation 35 from Section 7.3 of that paper into your existing logistic regression code. I know you already grasp the theory, so let's jump straight to mapping those paper terms to working Python code.

First, let's recap the key terms from Equation 35 to align with your setup:

Equation 35 (adapted for your logistic regression case):
$$
\mathcal{L}(\theta) = -\log p(y|X,\theta) + \lambda \sum_{k=1}^K \frac{1}{|D_k|} |\theta(D_k)|2^2 + \lambda |\theta{dyn}|_2^2
$$

  • $-\log p(y|X,\theta)$: The negative of your existing log_likelihood output (your current function returns the positive log-likelihood; we flip it to get a loss we can minimize)
  • $\lambda$: The regularization hyperparameter you'll tune later
  • $D_k$: Groups of static parameters (e.g., split your feature weights into logical groups)
  • $|D_k|$: Number of parameters in group $D_k$
  • $\theta(D_k)$: The subset of model weights belonging to group $D_k$
  • $\theta_{dyn}$: Dynamic parameters (typically the bias term, since it's not tied to a specific feature group)

Step 1: Update the Loss Function with Regularization

First, let's rewrite your log-likelihood function to return the regularized negative log-likelihood (the loss $\mathcal{L}$ we need to minimize):

def regularized_nll(features, target, weights, lambda_reg, param_groups, dyn_param_idx):
    # Calculate base negative log-likelihood (flip the sign of your original log-likelihood)
    scores = np.dot(features, weights)
    base_nll = -np.sum(target * scores - np.log(1 + np.exp(scores)))
    
    # Calculate static group regularization: λ * sum( (1/|D_k|) * ||θ(D_k)||² )
    reg_static = 0.0
    for group in param_groups:
        group_weights = weights[group]
        group_size = len(group)
        reg_static += (1 / group_size) * np.sum(group_weights ** 2)
    reg_static *= lambda_reg
    
    # Calculate dynamic parameter regularization: λ * ||θ_dyn||²
    dyn_weights = weights[dyn_param_idx]
    reg_dyn = lambda_reg * np.sum(dyn_weights ** 2)
    
    # Total loss is the sum of all terms
    total_loss = base_nll + reg_static + reg_dyn
    return total_loss

Parameter Breakdown:

  • lambda_reg: The $\lambda$ from the paper (start with 0.1 if you're testing)
  • param_groups: A list of index lists. For your simulated data, Xb has columns [bias, x1, x2], so you could use [[1], [2]] (each feature weight is its own group) or [[1,2]] (both features in one group)
  • dyn_param_idx: The index of your dynamic parameter (almost always the bias term, index 0 in your setup)

Step 2: Update Gradient Descent to Include Regularization

Next, we need to add the regularization terms to the gradient update. The gradient of the regularization terms is straightforward:

  • For each static group $D_k$: gradient = $\lambda * (2 / |D_k|) * \theta(D_k)$
  • For the dynamic parameter: gradient = $\lambda * 2 * \theta_{dyn}$

Here's the modified training function:

def regularized_log_reg(features, target, num_steps, learning_rate, lambda_reg, param_groups, dyn_param_idx):
    weights = np.zeros(features.shape[1])
    for step in range(num_steps):
        scores = np.dot(features, weights)
        predictions = sigmoid(scores)
        
        # Base gradient from standard logistic regression (maximizes log-likelihood)
        error = target - predictions
        gradient = np.dot(features.T, error)
        
        # Calculate regularization gradient (since we're minimizing loss, we subtract this from the base gradient)
        reg_gradient = np.zeros_like(weights)
        
        # Add static group regularization gradients
        for group in param_groups:
            group_weights = weights[group]
            group_size = len(group)
            reg_gradient[group] += (2 * lambda_reg / group_size) * group_weights
        
        # Add dynamic parameter regularization gradient
        reg_gradient[dyn_param_idx] += 2 * lambda_reg * weights[dyn_param_idx]
        
        # Update weights: adjust base gradient by regularization term
        gradient -= reg_gradient
        weights += learning_rate * gradient
        
        # Print progress every 10k steps
        if step % 10000 == 0:
            current_loss = regularized_nll(features, target, weights, lambda_reg, param_groups, dyn_param_idx)
            print(f"Step {step}, Total Loss: {current_loss:.4f}")
    return weights

Step 3: Test with Your Simulated Data

Let's run the regularized model with your existing simulated data. We'll use:

  • $\lambda = 0.1$ (you can tune this later with cross-validation)
  • Separate groups for each feature weight
  • Bias term as the dynamic parameter
# Regularization settings
lambda_reg = 0.1
param_groups = [[1], [2]]  # Feature 1 (index 1) = group 1, Feature 2 (index 2) = group 2
dyn_param_idx = [0]        # Bias term is the dynamic parameter

# Train the regularized model
regularized_weights = regularized_log_reg(
    Xb, simulated_labels, 
    num_steps=50000, learning_rate=5e-5,
    lambda_reg=lambda_reg, param_groups=param_groups, dyn_param_idx=dyn_param_idx
)

# Compare with unregularized weights (from your original function)
original_weights = log_reg(Xb, simulated_labels, num_steps=50000, learning_rate=5e-5)
print("\nOriginal Unregularized Weights:", original_weights)
print("Regularized Weights:", regularized_weights)

Quick Tips for Alignment with the Paper

  • If your parameter grouping is different (e.g., multiple features per group), just adjust the param_groups list to include the correct weight indices.
  • If you don't have dynamic parameters, simply remove the reg_dyn term from the loss and the corresponding gradient line.
  • Tune $\lambda$ using cross-validation (split your data into train/validation sets and pick the value that minimizes validation loss).

内容的提问来源于stack exchange,提问作者mkpisk

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.09 15:17:41