Python:如何利用CUDA将文件更快加载至内存?
Great question—accelerating dataset loading is such a critical bottleneck in most ML workflows, especially when you’re working with large-scale data. Since you’re already familiar with Python and CUDA, let’s dive into how you can leverage GPU power to speed this up beyond just relying on multiprocessing.
Traditionally, data travels from storage → CPU memory → GPU memory, which adds unnecessary copy overhead. With CUDA-based tools like GPUDirect Storage (GDS), you can cut out the CPU middleman entirely, transferring data directly from compatible storage (like NVMe SSDs) straight to GPU memory. This is a game-changer for large datasets where data movement often takes longer than model computation.
1. Use GPUDirect Storage (GDS) for Direct Disk-to-GPU Transfers
GDS is NVIDIA’s technology that enables direct data transfers between storage and GPU memory, bypassing CPU RAM entirely. In Python, you can use libraries like cupy (a GPU-accelerated drop-in replacement for NumPy) to leverage GDS:
import cupy as cp # Load a binary dataset directly into GPU memory with open("large_training_data.bin", "rb") as f: gpu_data = cp.load(f) # Now you can use gpu_data directly in CUDA-accelerated ML operations
Note: GDS requires compatible hardware (NVMe SSD, Volta or newer GPU) and proper driver/CUDA toolkit configuration.
2. Leverage ML Frameworks' GPU-Optimized Data Loaders
Most modern ML frameworks have built-in tools to streamline data loading to GPU, often working alongside multiprocessing for CPU-side preprocessing:
PyTorch Example
Use DataLoader with pin_memory=True to lock preprocessed data in CPU memory (reducing copy time to GPU) and num_workers for parallel CPU preprocessing:
from torch.utils.data import DataLoader, Dataset class CustomMLDataset(Dataset): def __getitem__(self, idx): # Load raw data and run CPU-side preprocessing (e.g., image decoding) raw_data = self.load_raw_data(idx) processed_data = self.preprocess(raw_data) # e.g., resize, normalize return processed_data # Configure loader for optimal GPU transfer data_loader = DataLoader( CustomMLDataset(), batch_size=64, num_workers=4, # Parallel CPU preprocessing pin_memory=True # Lock data in CPU RAM for faster GPU copy ) # Iterate and move batches directly to GPU for batch in data_loader: batch = batch.to("cuda") # Run model training/inference
Hugging Face Datasets Example
If you’re using Hugging Face’s datasets library, you can load data directly onto GPU with a single line:
from datasets import load_dataset # Load dataset and format it for GPU use dataset = load_dataset("cifar10", split="train") dataset.set_format("torch", device="cuda") # Access samples directly on GPU sample = dataset[0]["image"]
3. Use NVIDIA DALI for End-to-End GPU-Accelerated Data Pipelines
NVIDIA DALI (Data Loading Library) is purpose-built for GPU-accelerated data loading and preprocessing. It handles everything from reading files to applying augmentations entirely on the GPU, eliminating CPU bottlenecks:
from nvidia.dali.pipeline import Pipeline import nvidia.dali.fn as fn import nvidia.dali.types as types # Define a DALI pipeline for image data pipe = Pipeline(batch_size=32, num_threads=4, device_id=0) with pipe: # Read files from disk jpegs, labels = fn.readers.file(file_root="path/to/imagenet/train", random_shuffle=True) # Decode images directly on GPU (or mixed CPU/GPU) images = fn.decoders.image(jpegs, device="mixed", output_type=types.RGB) # Apply GPU-accelerated augmentations images = fn.resize(images, resize_x=224, resize_y=224) images = fn.normalize(images, mean=[0.485, 0.456, 0.406], std=[0.229, 0.224, 0.225]) pipe.set_outputs(images, labels) # Build and run the pipeline pipe.build() for _ in range(pipe.epoch_size() // 32): gpu_images, gpu_labels = pipe.run() # Convert to PyTorch/TensorFlow tensors if needed torch_images = gpu_images.as_tensor()
DALI integrates seamlessly with PyTorch, TensorFlow, and MXNet, making it easy to drop into existing workflows.
4. Optimize Data Formats for GPU Loading
Binary formats like HDF5, TFRecord, or Parquet load much faster than text-based formats (like CSV) because they require less parsing. Pre-convert your dataset to one of these formats to reduce initial loading time, then pair it with GPU-accelerated readers (like cupy for HDF5, or TensorFlow’s TFRecord reader).
- Hardware Limits: GPU memory is often smaller than CPU RAM, so you’ll need to batch data carefully or use mixed precision to fit large datasets.
- Overhead: For small datasets, CPU loading might still be faster—GPU initialization and transfer overhead can outweigh benefits for tiny batches.
- Compatibility: Ensure your storage (NVMe SSD) and GPU (Volta+) support GDS if you’re using that method.
Hope these tips help you slash your dataset loading times! It’s all about matching the right tool to your workflow and hardware.
内容的提问来源于stack exchange,提问作者Surya

