PyTorch分布式多GPU推理输出一致问题求助
Great question! Let's break down why you're seeing identical acc1 values across all 4 GPUs, even though you expect each to process different batches.
Most Likely Cause: Incorrect Validation Data Sampling
The biggest red flag here is that all GPUs are processing the exact same first batch of data. This happens when you don't use a DistributedSampler for your validation loader.
By default, if you use a regular SequentialSampler (or no sampler at all) in a distributed setup, each GPU process will load the full validation dataset independently. That means every GPU starts with the first 32 images of the dataset, hence identical outputs and accuracy for the first batch (i=0 in your loop).
How to Fix It
You need to explicitly use torch.utils.data.distributed.DistributedSampler for your validation dataset to ensure each GPU gets a unique, non-overlapping slice of the data. Here's how to adjust your data loader setup:
from torch.utils.data.distributed import DistributedSampler # Initialize validation dataset (your existing code here) val_dataset = ... # Create distributed sampler for validation (shuffle=False since we're evaluating) val_sampler = DistributedSampler(val_dataset, shuffle=False) # Create validation loader with the sampler val_loader = torch.utils.data.DataLoader( val_dataset, batch_size=args.batch_size // args.world_size, # 32 per GPU when total batch size is 128 sampler=val_sampler, num_workers=args.workers, pin_memory=True, non_blocking=True )
Also, make sure to call val_sampler.set_epoch(0) before starting the validation loop (even though we're not shuffling, this is a good practice to ensure consistency across runs).
Other Minor Notes
- You have a duplicate line in your code:
top1.update(acc1[0], images.size(0))is called twice. This won't cause the identical accuracy issue, but it will incorrectly double-count the first batch's metrics—worth fixing! - Double-check that your checkpoint is loaded correctly on all GPUs. While this is less likely to cause identical accuracies for the first batch, ensuring all GPUs have the same model weights is foundational for distributed evaluation.
Verify the Fix
After updating the sampler, re-run your code and check the first batch outputs. Each GPU should now process a different subset of 32 images, leading to distinct acc1 values.
内容的提问来源于stack exchange,提问作者WButter

