如何优化Numpy实现Skip-Gram时大型向量列表的追加效率?
解决Skip-Gram独热向量内存占用与耗时问题
你的核心问题在于完全没必要存储完整的独热向量——独热向量是极端稀疏的,只有对应词索引的位置为1,其余全为0,存储全量的0元素纯粹是内存和时间的浪费。以下是纯Python+Numpy的高效解决方案:
1. 存储跳元对的词索引而非独热向量
直接收集目标词与上下文词的整数索引,而非生成50000维的全零向量。这种方式内存占用微乎其微,生成速度极快:
import numpy as np from tqdm import tqdm vocab_size = 50000 num_sentences = 100000 avg_words_per_sentence = 15 # 模拟生成句子的词索引(替换为你的真实数据逻辑) def get_sentence_indices(): # 生成长度在10-20之间的句子索引 return np.random.randint(0, vocab_size, size=np.random.randint(10, 20)) xy = [] window_size = 2 # Skip-Gram的上下文窗口大小,可按需调整 for _ in tqdm(range(num_sentences)): word_indices = get_sentence_indices() sentence_len = len(word_indices) # 遍历每个目标词,生成对应的上下文跳元对 for target_pos in range(sentence_len): target_idx = word_indices[target_pos] # 确定上下文的范围(避免越界) start = max(0, target_pos - window_size) end = min(sentence_len, target_pos + window_size + 1) # 跳过目标词自身,收集上下文词索引 for context_pos in range(start, end): if context_pos != target_pos: context_idx = word_indices[context_pos] xy.append((target_idx, context_idx))
这段代码生成的xy是[(目标词索引, 上下文词索引), ...]的列表,每个元素仅占用两个整数的内存,150万条数据仅需约12MB内存,完全不会出现内存溢出。
2. 训练时直接用索引做矩阵运算
Skip-Gram的核心计算不需要依赖独热向量:
- 目标词的独热向量与输入权重矩阵
W(形状为(vocab_size, embed_dim))的乘积,等价于直接取W的第target_idx行; - 上下文词的预测概率计算,也可通过索引直接定位到对应位置的得分。
以下是核心训练逻辑示例:
embed_dim = 100 # 词向量维度,可调整 # 初始化权重矩阵 W = np.random.randn(vocab_size, embed_dim) * 0.01 # 输入层到隐藏层权重 W_prime = np.random.randn(embed_dim, vocab_size) * 0.01 # 隐藏层到输出层权重 batch_size = 256 num_batches = len(xy) // batch_size learning_rate = 0.001 for batch_idx in tqdm(range(num_batches)): # 获取当前批量的索引对 batch_pairs = xy[batch_idx*batch_size : (batch_idx+1)*batch_size] t_indices = np.array([p[0] for p in batch_pairs]) c_indices = np.array([p[1] for p in batch_pairs]) # 批量计算隐藏层向量(等价于独热向量与W的乘积) hidden = W[t_indices] # 形状: (batch_size, embed_dim) # 计算输出层得分 output = hidden.dot(W_prime) # 形状: (batch_size, vocab_size) # 数值稳定的批量Softmax计算 max_output = np.max(output, axis=1, keepdims=True) exp_output = np.exp(output - max_output) prob = exp_output / np.sum(exp_output, axis=1, keepdims=True) # 计算批量交叉熵损失 batch_loss = -np.log(prob[np.arange(batch_size), c_indices]) avg_loss = np.mean(batch_loss) # 梯度计算与更新(Skip-Gram标准梯度) grad_output = prob.copy() grad_output[np.arange(batch_size), c_indices] -= 1 grad_W_prime = hidden.T.dot(grad_output) grad_W = grad_output.dot(W_prime.T) # 更新权重 W_prime -= learning_rate * grad_W_prime / batch_size W[t_indices] -= learning_rate * grad_W / batch_size
3. 额外优化建议
- 打乱数据:在训练前打乱
xy列表,避免模型陷入局部最优; - 使用负采样:如果vocab_size较大(50000属于较大规模),全量计算Softmax耗时仍会很高,可实现Skip-Gram的负采样版本,仅计算正样本和少量负样本的损失,进一步降低计算量;
- 内存复用:避免在循环中重复创建新数组,可预先分配批量索引的数组空间。
内容的提问来源于stack exchange,提问作者Mahesha999
相关产品推荐
相关产品推荐

