You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何查看HuggingFace模型训练后的参数变化?——BART摘要模型微调后参数验证困惑解答

How to Identify Which Model Parameters Were Modified During BART Fine-Tuning

Great question! The confusion here boils down to a critical distinction: model configuration defines the fixed architecture (like layer count, hidden size, attention heads) while trainable weights are what actually update during fine-tuning. The model.config and print(model) output only show structural details, not the underlying weight values—that’s why they look identical pre- and post-fine-tuning, even though your model generates different summaries.

Here are practical, actionable steps to pinpoint exactly which parameters changed:


1. Directly Compare Weights Using state_dict()

Every PyTorch model stores its trainable (and non-trainable) weights in a state_dict()—a dictionary mapping parameter names to their tensor values. Use this to compare your pre- and post-fine-tuning models:

import torch

# Extract weight dictionaries for both models
state_dict_before = model_before_tuning_1.state_dict()
state_dict_after = model.state_dict()

# Track parameters that changed
modified_params = []
for param_name in state_dict_before.keys():
    # Skip non-trainable parameters (if any were frozen)
    if not next(p for n, p in model.named_parameters() if n == param_name).requires_grad:
        continue
    # Check if tensors are numerically different (use atol for floating-point stability)
    if not torch.allclose(state_dict_before[param_name], state_dict_after[param_name], atol=1e-6):
        modified_params.append(param_name)

print(f"Total modified parameters: {len(modified_params)}")
print("Sample modified parameters:", modified_params[:10])  # Show first 10 examples

For BART, you’ll see parameter names like:

  • encoder.layers.3.self_attn.v_proj.weight
  • decoder.layers.1.fc2.bias
  • lm_head.weight (the final generation head, shared with input embeddings)

2. Focus Only on Trainable Parameters

If you froze some layers during fine-tuning (even if you didn’t, this filters out non-trainable metadata), use this to narrow down to parameters that could have changed:

# Get names of parameters marked as trainable in your fine-tuned model
trainable_names = [name for name, param in model.named_parameters() if param.requires_grad]

# Compare only these parameters
modified_trainable = []
for name in trainable_names:
    if not torch.allclose(state_dict_before[name], state_dict_after[name], atol=1e-6):
        modified_trainable.append(name)

3. Quantify and Visualize Parameter Changes

To understand how much each parameter shifted, calculate metrics like the L2 norm of the weight difference:

param_change_magnitude = {}
for name in modified_trainable:
    weight_diff = state_dict_after[name] - state_dict_before[name]
    l2_norm = torch.norm(weight_diff).item()
    param_change_magnitude[name] = l2_norm

# Sort parameters by largest change
sorted_changes = sorted(param_change_magnitude.items(), key=lambda x: x[1], reverse=True)
print("Parameters with the largest updates:")
for name, norm in sorted_changes[:5]:
    print(f"{name}: L2 Norm = {norm:.4f}")

You can also visualize these changes with histograms or heatmaps (using Matplotlib/Seaborn) to spot patterns—for example, if decoder layers changed more than encoder layers, or specific attention heads had bigger updates.


Why model.config and print(model) Stay Identical

Just to clarify:

  • model.config stores architectural metadata (layer count, hidden dimension, dropout rates, etc.). Fine-tuning never alters the model’s structure, so this remains unchanged.
  • print(model) outputs the model’s class hierarchy and layer structure, but not the actual weight values. Two models with the same architecture will look identical here, even if their weights are completely different.

内容的提问来源于stack exchange,提问作者Zenith_Raven

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.01 02:02:46