You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Huggingface Trainer批量微调模型时Wandb仅展示首个模型的图表与日志问题

Solution: Wandb Only Shows First Model's Logs in Batch Finetuning

The issue you're facing is super common—it happens because Wandb reuses the same run session by default when you call finetune() in a loop. Subsequent training tasks don't spin up new Wandb runs, so all logs get merged into the first one. That's why you only see results for the first file, even though all models finish training and save correctly. Here's how to fix this:


1. Explicitly Manage Wandb Runs Per Training Task

You need to initialize a unique Wandb run for each file and close it after training wraps up. This ensures every finetuning job gets its own separate log entry in Wandb.

Modify your finetune function like this:

import wandb

def finetune(args, file):
    # Initialize a unique Wandb run for this specific file
    wandb.init(
        project="your-wandb-project-name",  # Replace with your actual project name
        name=f"{model_name}-finetuned-{file}",  # Unique name to distinguish runs
        config=vars(args)  # Optional: Sync your training arguments to Wandb for tracking
    )

    training_args = TrainingArguments(
        output_dir=f'{model_name}-finetuned-{file}',
        overwrite_output_dir=True,
        evaluation_strategy='no',
        num_train_epochs=args.epochs,
        learning_rate=args.lr,
        weight_decay=args.decay,
        per_device_train_batch_size=args.batch_size,
        per_device_eval_batch_size=args.batch_size,
        fp16=True,
        save_strategy='no',
        seed=args.seed,
        dataloader_num_workers=4,
        # Tell the Trainer to send logs to Wandb
        report_to="wandb",
        # Match the run name to keep logs consistent with your model output
        run_name=f"{model_name}-finetuned-{file}"
    )

    trainer = Trainer(
        model=model,
        args=training_args,
        train_dataset=tokenized_dataset['train'],
        eval_dataset=None,
        data_collator=data_collator,
    )
    trainer.train()
    trainer.save_model()

    # Close the current Wandb run to prepare for the next finetuning task
    wandb.finish()

2. Load a Fresh Model Instance Every Time

If you're reusing the same model object across all finetuning tasks (loaded once outside the loop), you're not starting fresh—you're continuing training on the same weights from the previous file. This will mess up your model performance and make Wandb logs even more inconsistent.

Add model reloading inside the finetune function:

from transformers import AutoModelForSequenceClassification  # Adjust based on your task type

def finetune(args, file):
    # Load a brand new pre-trained model for each finetuning job
    model = AutoModelForSequenceClassification.from_pretrained(
        model_name,
        num_labels=args.num_labels  # Replace with your task's number of labels
    )

    # Rest of the Wandb init and training code goes here...

3. Double-Check Wandb Sync Settings

Make sure report_to="wandb" is set in your TrainingArguments—this is what tells the Hugging Face Trainer to automatically sync metrics and logs to Wandb. Without this line, subsequent tasks might not log any data at all.

After making these changes, each file's finetuning job will create a separate run in Wandb, and you'll be able to view all your models' charts, metrics, and logs individually.

内容的提问来源于stack exchange,提问作者kkgarg

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.28 18:42:30