You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

能否在GPU上运行含Numpy的常规Python循环?附示例代码

能否在GPU上运行带Numpy操作的Python循环?

当然可以,但原生Numpy和常规Python循环本身是跑在CPU上的——得做一些针对性调整才能利用GPU算力。结合你提供的字符级数据预处理代码(处理1800万+字符、394词汇量、seq_length=30),给你几个实用方案:

方案1:用CuPy替换Numpy(改动最小)

CuPy是NVIDIA推出的GPU版Numpy,API和Numpy几乎完全一致,只需替换导入语句,就能让数组操作自动跑在GPU上。虽然Python的for循环本身还是在CPU执行,但循环内的核心计算会被GPU加速,对于你的大数据量来说能明显提升速度。

修改后的代码示例:

import cupy as cp  # 替换import numpy as np

def data_preprocess(data_dir, seq_length):
    with open(data_dir, 'r', encoding="utf8") as f:
        data = f.read()
    
    chars = sorted(list(set(data)))
    VOCAB_SIZE = len(chars)
    print('Data length: {} characters'.format(len(data)))
    print('Vocabulary size: {} characters'.format(VOCAB_SIZE))
    
    ix_to_char = {ix: char for ix, char in enumerate(chars)}
    char_to_ix = {char: ix for ix, char in enumerate(chars)}
    
    total_sequences = len(data) // seq_length
    X = cp.zeros((total_sequences, seq_length, VOCAB_SIZE))
    y = cp.zeros((total_sequences, seq_length, VOCAB_SIZE))
    
    for i in range(total_sequences):
        # 处理输入序列
        X_sequence = data[i * seq_length:(i + 1) * seq_length]
        X_sequence_ix = [char_to_ix[value] for value in X_sequence]
        # CuPy的数组操作自动跑在GPU上
        input_sequence = cp.zeros((seq_length, VOCAB_SIZE))
        for j in range(seq_length):
            input_sequence[j][X_sequence_ix[j]] = 1.
        X[i] = input_sequence
        
        # 处理目标序列
        y_sequence = data[i * seq_length + 1:(i + 1) * seq_length + 1]
        y_sequence_ix = [char_to_ix[value] for value in y_sequence]
        target_sequence = cp.zeros((seq_length, VOCAB_SIZE))
        for j in range(seq_length):
            target_sequence[j][y_sequence_ix[j]] = 1.
        y[i] = target_sequence
    
    # 如果后续需要转回CPU的Numpy数组,调用X.get()、y.get()即可
    return X, y, VOCAB_SIZE, ix_to_char

方案2:向量化操作+CuPy(大幅提升性能)

你的代码里有多层嵌套for循环,这是性能瓶颈——哪怕用了CuPy,Python循环本身的开销还是存在。可以把字符转索引、one-hot编码的逻辑改成向量化操作,彻底去掉内层循环:

import cupy as cp

def data_preprocess(data_dir, seq_length):
    with open(data_dir, 'r', encoding="utf8") as f:
        data = f.read()
    
    chars = sorted(list(set(data)))
    VOCAB_SIZE = len(chars)
    print('Data length: {} characters'.format(len(data)))
    print('Vocabulary size: {} characters'.format(VOCAB_SIZE))
    
    ix_to_char = {ix: char for ix, char in enumerate(chars)}
    char_to_ix = {char: ix for ix, char in enumerate(chars)}
    
    # 把整个数据集转成GPU上的索引数组
    data_ix = cp.array([char_to_ix[c] for c in data])
    total_sequences = len(data) // seq_length
    
    # 切分输入和目标序列的索引(向量化切分,替代外层循环的部分逻辑)
    X_ix = data_ix[:total_sequences * seq_length].reshape(total_sequences, seq_length)
    y_ix = data_ix[1:total_sequences * seq_length + 1].reshape(total_sequences, seq_length)
    
    # 用CuPy的eye函数向量化生成one-hot编码,彻底去掉内层循环
    X = cp.eye(VOCAB_SIZE)[X_ix]
    y = cp.eye(VOCAB_SIZE)[y_ix]
    
    return X, y, VOCAB_SIZE, ix_to_char

这个版本把原本的两层循环压缩成了向量化操作,GPU的并行计算能力能被充分利用,处理1800万字符的速度会比原代码快一个数量级以上。

方案3:用深度学习框架(TensorFlow/PyTorch)——最适合后续模型训练

既然你在做序列数据预处理,大概率是为了训练RNN/Transformer这类模型。直接用TensorFlow或PyTorch的GPU张量来处理,不仅能实现GPU加速,还能无缝衔接后续的模型训练流程,避免数据在CPU和GPU之间来回拷贝的开销。

以PyTorch为例的代码:

import torch

def data_preprocess(data_dir, seq_length):
    with open(data_dir, 'r', encoding="utf8") as f:
        data = f.read()
    
    chars = sorted(list(set(data)))
    VOCAB_SIZE = len(chars)
    print('Data length: {} characters'.format(len(data)))
    print('Vocabulary size: {} characters'.format(VOCAB_SIZE))
    
    ix_to_char = {ix: char for ix, char in enumerate(chars)}
    char_to_ix = {char: ix for ix, char in enumerate(chars)}
    
    # 自动检测GPU设备
    device = torch.device('cuda' if torch.cuda.is_available() else 'cpu')
    # 把数据转成GPU上的长整型张量
    data_ix = torch.tensor([char_to_ix[c] for c in data], dtype=torch.long).to(device)
    total_sequences = len(data) // seq_length
    
    # 向量化切分序列
    X_ix = data_ix[:total_sequences * seq_length].view(total_sequences, seq_length)
    y_ix = data_ix[1:total_sequences * seq_length + 1].view(total_sequences, seq_length)
    
    # 用PyTorch内置函数生成one-hot编码
    X = torch.nn.functional.one_hot(X_ix, num_classes=VOCAB_SIZE).float()
    y = torch.nn.functional.one_hot(y_ix, num_classes=VOCAB_SIZE).float()
    
    return X, y, VOCAB_SIZE, ix_to_char

总结

  • 原生Numpy+Python循环无法利用GPU,必须替换为GPU加速的库或框架;
  • 追求最小改动选CuPy;
  • 追求最高性能选「CuPy+向量化操作」;
  • 如果后续要训练深度学习模型,直接用TensorFlow/PyTorch是最优解。

内容的提问来源于stack exchange,提问作者user7935479

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.29 08:05:58