能否在GPU上运行含Numpy的常规Python循环?附示例代码
能否在GPU上运行带Numpy操作的Python循环?
当然可以,但原生Numpy和常规Python循环本身是跑在CPU上的——得做一些针对性调整才能利用GPU算力。结合你提供的字符级数据预处理代码(处理1800万+字符、394词汇量、seq_length=30),给你几个实用方案:
方案1:用CuPy替换Numpy(改动最小)
CuPy是NVIDIA推出的GPU版Numpy,API和Numpy几乎完全一致,只需替换导入语句,就能让数组操作自动跑在GPU上。虽然Python的for循环本身还是在CPU执行,但循环内的核心计算会被GPU加速,对于你的大数据量来说能明显提升速度。
修改后的代码示例:
import cupy as cp # 替换import numpy as np def data_preprocess(data_dir, seq_length): with open(data_dir, 'r', encoding="utf8") as f: data = f.read() chars = sorted(list(set(data))) VOCAB_SIZE = len(chars) print('Data length: {} characters'.format(len(data))) print('Vocabulary size: {} characters'.format(VOCAB_SIZE)) ix_to_char = {ix: char for ix, char in enumerate(chars)} char_to_ix = {char: ix for ix, char in enumerate(chars)} total_sequences = len(data) // seq_length X = cp.zeros((total_sequences, seq_length, VOCAB_SIZE)) y = cp.zeros((total_sequences, seq_length, VOCAB_SIZE)) for i in range(total_sequences): # 处理输入序列 X_sequence = data[i * seq_length:(i + 1) * seq_length] X_sequence_ix = [char_to_ix[value] for value in X_sequence] # CuPy的数组操作自动跑在GPU上 input_sequence = cp.zeros((seq_length, VOCAB_SIZE)) for j in range(seq_length): input_sequence[j][X_sequence_ix[j]] = 1. X[i] = input_sequence # 处理目标序列 y_sequence = data[i * seq_length + 1:(i + 1) * seq_length + 1] y_sequence_ix = [char_to_ix[value] for value in y_sequence] target_sequence = cp.zeros((seq_length, VOCAB_SIZE)) for j in range(seq_length): target_sequence[j][y_sequence_ix[j]] = 1. y[i] = target_sequence # 如果后续需要转回CPU的Numpy数组,调用X.get()、y.get()即可 return X, y, VOCAB_SIZE, ix_to_char
方案2:向量化操作+CuPy(大幅提升性能)
你的代码里有多层嵌套for循环,这是性能瓶颈——哪怕用了CuPy,Python循环本身的开销还是存在。可以把字符转索引、one-hot编码的逻辑改成向量化操作,彻底去掉内层循环:
import cupy as cp def data_preprocess(data_dir, seq_length): with open(data_dir, 'r', encoding="utf8") as f: data = f.read() chars = sorted(list(set(data))) VOCAB_SIZE = len(chars) print('Data length: {} characters'.format(len(data))) print('Vocabulary size: {} characters'.format(VOCAB_SIZE)) ix_to_char = {ix: char for ix, char in enumerate(chars)} char_to_ix = {char: ix for ix, char in enumerate(chars)} # 把整个数据集转成GPU上的索引数组 data_ix = cp.array([char_to_ix[c] for c in data]) total_sequences = len(data) // seq_length # 切分输入和目标序列的索引(向量化切分,替代外层循环的部分逻辑) X_ix = data_ix[:total_sequences * seq_length].reshape(total_sequences, seq_length) y_ix = data_ix[1:total_sequences * seq_length + 1].reshape(total_sequences, seq_length) # 用CuPy的eye函数向量化生成one-hot编码,彻底去掉内层循环 X = cp.eye(VOCAB_SIZE)[X_ix] y = cp.eye(VOCAB_SIZE)[y_ix] return X, y, VOCAB_SIZE, ix_to_char
这个版本把原本的两层循环压缩成了向量化操作,GPU的并行计算能力能被充分利用,处理1800万字符的速度会比原代码快一个数量级以上。
方案3:用深度学习框架(TensorFlow/PyTorch)——最适合后续模型训练
既然你在做序列数据预处理,大概率是为了训练RNN/Transformer这类模型。直接用TensorFlow或PyTorch的GPU张量来处理,不仅能实现GPU加速,还能无缝衔接后续的模型训练流程,避免数据在CPU和GPU之间来回拷贝的开销。
以PyTorch为例的代码:
import torch def data_preprocess(data_dir, seq_length): with open(data_dir, 'r', encoding="utf8") as f: data = f.read() chars = sorted(list(set(data))) VOCAB_SIZE = len(chars) print('Data length: {} characters'.format(len(data))) print('Vocabulary size: {} characters'.format(VOCAB_SIZE)) ix_to_char = {ix: char for ix, char in enumerate(chars)} char_to_ix = {char: ix for ix, char in enumerate(chars)} # 自动检测GPU设备 device = torch.device('cuda' if torch.cuda.is_available() else 'cpu') # 把数据转成GPU上的长整型张量 data_ix = torch.tensor([char_to_ix[c] for c in data], dtype=torch.long).to(device) total_sequences = len(data) // seq_length # 向量化切分序列 X_ix = data_ix[:total_sequences * seq_length].view(total_sequences, seq_length) y_ix = data_ix[1:total_sequences * seq_length + 1].view(total_sequences, seq_length) # 用PyTorch内置函数生成one-hot编码 X = torch.nn.functional.one_hot(X_ix, num_classes=VOCAB_SIZE).float() y = torch.nn.functional.one_hot(y_ix, num_classes=VOCAB_SIZE).float() return X, y, VOCAB_SIZE, ix_to_char
总结
- 原生Numpy+Python循环无法利用GPU,必须替换为GPU加速的库或框架;
- 追求最小改动选CuPy;
- 追求最高性能选「CuPy+向量化操作」;
- 如果后续要训练深度学习模型,直接用TensorFlow/PyTorch是最优解。
内容的提问来源于stack exchange,提问作者user7935479
相关产品推荐
相关产品推荐

