You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何优化Numpy实现Skip-Gram时大型向量列表的追加效率?

解决Skip-Gram独热向量内存占用与耗时问题

你的核心问题在于完全没必要存储完整的独热向量——独热向量是极端稀疏的,只有对应词索引的位置为1,其余全为0,存储全量的0元素纯粹是内存和时间的浪费。以下是纯Python+Numpy的高效解决方案:

1. 存储跳元对的词索引而非独热向量

直接收集目标词与上下文词的整数索引,而非生成50000维的全零向量。这种方式内存占用微乎其微,生成速度极快:

import numpy as np
from tqdm import tqdm

vocab_size = 50000
num_sentences = 100000
avg_words_per_sentence = 15

# 模拟生成句子的词索引(替换为你的真实数据逻辑)
def get_sentence_indices():
    # 生成长度在10-20之间的句子索引
    return np.random.randint(0, vocab_size, size=np.random.randint(10, 20))

xy = []
window_size = 2  # Skip-Gram的上下文窗口大小,可按需调整
for _ in tqdm(range(num_sentences)):
    word_indices = get_sentence_indices()
    sentence_len = len(word_indices)
    # 遍历每个目标词,生成对应的上下文跳元对
    for target_pos in range(sentence_len):
        target_idx = word_indices[target_pos]
        # 确定上下文的范围(避免越界)
        start = max(0, target_pos - window_size)
        end = min(sentence_len, target_pos + window_size + 1)
        # 跳过目标词自身,收集上下文词索引
        for context_pos in range(start, end):
            if context_pos != target_pos:
                context_idx = word_indices[context_pos]
                xy.append((target_idx, context_idx))

这段代码生成的xy是[(目标词索引, 上下文词索引), ...]的列表,每个元素仅占用两个整数的内存,150万条数据仅需约12MB内存,完全不会出现内存溢出。

2. 训练时直接用索引做矩阵运算

Skip-Gram的核心计算不需要依赖独热向量:

  • 目标词的独热向量与输入权重矩阵W(形状为(vocab_size, embed_dim))的乘积,等价于直接取W的第target_idx行;
  • 上下文词的预测概率计算,也可通过索引直接定位到对应位置的得分。

以下是核心训练逻辑示例:

embed_dim = 100  # 词向量维度,可调整
# 初始化权重矩阵
W = np.random.randn(vocab_size, embed_dim) * 0.01  # 输入层到隐藏层权重
W_prime = np.random.randn(embed_dim, vocab_size) * 0.01  # 隐藏层到输出层权重

batch_size = 256
num_batches = len(xy) // batch_size
learning_rate = 0.001

for batch_idx in tqdm(range(num_batches)):
    # 获取当前批量的索引对
    batch_pairs = xy[batch_idx*batch_size : (batch_idx+1)*batch_size]
    t_indices = np.array([p[0] for p in batch_pairs])
    c_indices = np.array([p[1] for p in batch_pairs])
    
    # 批量计算隐藏层向量(等价于独热向量与W的乘积)
    hidden = W[t_indices]  # 形状: (batch_size, embed_dim)
    # 计算输出层得分
    output = hidden.dot(W_prime)  # 形状: (batch_size, vocab_size)
    
    # 数值稳定的批量Softmax计算
    max_output = np.max(output, axis=1, keepdims=True)
    exp_output = np.exp(output - max_output)
    prob = exp_output / np.sum(exp_output, axis=1, keepdims=True)
    
    # 计算批量交叉熵损失
    batch_loss = -np.log(prob[np.arange(batch_size), c_indices])
    avg_loss = np.mean(batch_loss)
    
    # 梯度计算与更新(Skip-Gram标准梯度)
    grad_output = prob.copy()
    grad_output[np.arange(batch_size), c_indices] -= 1
    grad_W_prime = hidden.T.dot(grad_output)
    grad_W = grad_output.dot(W_prime.T)
    
    # 更新权重
    W_prime -= learning_rate * grad_W_prime / batch_size
    W[t_indices] -= learning_rate * grad_W / batch_size

3. 额外优化建议

  • 打乱数据:在训练前打乱xy列表,避免模型陷入局部最优;
  • 使用负采样:如果vocab_size较大(50000属于较大规模),全量计算Softmax耗时仍会很高,可实现Skip-Gram的负采样版本,仅计算正样本和少量负样本的损失,进一步降低计算量;
  • 内存复用:避免在循环中重复创建新数组,可预先分配批量索引的数组空间。

内容的提问来源于stack exchange,提问作者Mahesha999

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.01 12:45:35