You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PyTorch分布式多GPU推理输出一致问题求助

Why All GPUs Show the Same Accuracy in PyTorch Distributed Evaluation?

Great question! Let's break down why you're seeing identical acc1 values across all 4 GPUs, even though you expect each to process different batches.

Most Likely Cause: Incorrect Validation Data Sampling

The biggest red flag here is that all GPUs are processing the exact same first batch of data. This happens when you don't use a DistributedSampler for your validation loader.

By default, if you use a regular SequentialSampler (or no sampler at all) in a distributed setup, each GPU process will load the full validation dataset independently. That means every GPU starts with the first 32 images of the dataset, hence identical outputs and accuracy for the first batch (i=0 in your loop).

How to Fix It

You need to explicitly use torch.utils.data.distributed.DistributedSampler for your validation dataset to ensure each GPU gets a unique, non-overlapping slice of the data. Here's how to adjust your data loader setup:

from torch.utils.data.distributed import DistributedSampler

# Initialize validation dataset (your existing code here)
val_dataset = ... 

# Create distributed sampler for validation (shuffle=False since we're evaluating)
val_sampler = DistributedSampler(val_dataset, shuffle=False)

# Create validation loader with the sampler
val_loader = torch.utils.data.DataLoader(
    val_dataset,
    batch_size=args.batch_size // args.world_size,  # 32 per GPU when total batch size is 128
    sampler=val_sampler,
    num_workers=args.workers,
    pin_memory=True,
    non_blocking=True
)

Also, make sure to call val_sampler.set_epoch(0) before starting the validation loop (even though we're not shuffling, this is a good practice to ensure consistency across runs).

Other Minor Notes

  • You have a duplicate line in your code: top1.update(acc1[0], images.size(0)) is called twice. This won't cause the identical accuracy issue, but it will incorrectly double-count the first batch's metrics—worth fixing!
  • Double-check that your checkpoint is loaded correctly on all GPUs. While this is less likely to cause identical accuracies for the first batch, ensuring all GPUs have the same model weights is foundational for distributed evaluation.

Verify the Fix

After updating the sampler, re-run your code and check the first batch outputs. Each GPU should now process a different subset of 32 images, leading to distinct acc1 values.

内容的提问来源于stack exchange,提问作者WButter

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 06:46:46