如何查看HuggingFace模型训练后的参数变化?——BART摘要模型微调后参数验证困惑解答
Great question! The confusion here boils down to a critical distinction: model configuration defines the fixed architecture (like layer count, hidden size, attention heads) while trainable weights are what actually update during fine-tuning. The model.config and print(model) output only show structural details, not the underlying weight values—that’s why they look identical pre- and post-fine-tuning, even though your model generates different summaries.
Here are practical, actionable steps to pinpoint exactly which parameters changed:
1. Directly Compare Weights Using state_dict()
Every PyTorch model stores its trainable (and non-trainable) weights in a state_dict()—a dictionary mapping parameter names to their tensor values. Use this to compare your pre- and post-fine-tuning models:
import torch # Extract weight dictionaries for both models state_dict_before = model_before_tuning_1.state_dict() state_dict_after = model.state_dict() # Track parameters that changed modified_params = [] for param_name in state_dict_before.keys(): # Skip non-trainable parameters (if any were frozen) if not next(p for n, p in model.named_parameters() if n == param_name).requires_grad: continue # Check if tensors are numerically different (use atol for floating-point stability) if not torch.allclose(state_dict_before[param_name], state_dict_after[param_name], atol=1e-6): modified_params.append(param_name) print(f"Total modified parameters: {len(modified_params)}") print("Sample modified parameters:", modified_params[:10]) # Show first 10 examples
For BART, you’ll see parameter names like:
encoder.layers.3.self_attn.v_proj.weightdecoder.layers.1.fc2.biaslm_head.weight(the final generation head, shared with input embeddings)
2. Focus Only on Trainable Parameters
If you froze some layers during fine-tuning (even if you didn’t, this filters out non-trainable metadata), use this to narrow down to parameters that could have changed:
# Get names of parameters marked as trainable in your fine-tuned model trainable_names = [name for name, param in model.named_parameters() if param.requires_grad] # Compare only these parameters modified_trainable = [] for name in trainable_names: if not torch.allclose(state_dict_before[name], state_dict_after[name], atol=1e-6): modified_trainable.append(name)
3. Quantify and Visualize Parameter Changes
To understand how much each parameter shifted, calculate metrics like the L2 norm of the weight difference:
param_change_magnitude = {} for name in modified_trainable: weight_diff = state_dict_after[name] - state_dict_before[name] l2_norm = torch.norm(weight_diff).item() param_change_magnitude[name] = l2_norm # Sort parameters by largest change sorted_changes = sorted(param_change_magnitude.items(), key=lambda x: x[1], reverse=True) print("Parameters with the largest updates:") for name, norm in sorted_changes[:5]: print(f"{name}: L2 Norm = {norm:.4f}")
You can also visualize these changes with histograms or heatmaps (using Matplotlib/Seaborn) to spot patterns—for example, if decoder layers changed more than encoder layers, or specific attention heads had bigger updates.
Why model.config and print(model) Stay Identical
Just to clarify:
model.configstores architectural metadata (layer count, hidden dimension, dropout rates, etc.). Fine-tuning never alters the model’s structure, so this remains unchanged.print(model)outputs the model’s class hierarchy and layer structure, but not the actual weight values. Two models with the same architecture will look identical here, even if their weights are completely different.
内容的提问来源于stack exchange,提问作者Zenith_Raven

