You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

内存受限下,如何计算向量点积并检索最相似训练集ID?

内存受限下的大规模向量相似度匹配解决方案

问题背景

我需要计算测试文件与多份训练文件中句子的向量点积,以获取最相似的ID:

测试文件信息

from datasets import load_dataset
test_file = load_dataset('parquet', data_files='test_file.parquet')['train']
test_file
>>> features: ['sentences'] # 长度约10万条

训练文件信息

for idx in range(30):
    training_list_chunk = load_dataset('parquet', data_files=f'file_{idx}.parquet')['train']  # 长度约10万条
>>> features: ['sentences', 'ids'] # 每个文件含句子字符串列表和范围[0,30M]的整数ID列

训练集总计约3000万条数据(30个文件各10万条)。我使用Contriever模型将句子转换为向量,代码如下:

from transformers import AutoTokenizer, AutoModel
import torch

device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
tokenizer = AutoTokenizer.from_pretrained('facebook/contriever')
model = AutoModel.from_pretrained('facebook/contriever').to(device)

# 编码句子
sentences_inputs = tokenizer(list_of_strings, padding=True, truncation=True, return_tensors='pt')

# 计算token嵌入
outputs = model(**sentences_inputs.to(device))

# 均值池化函数
def mean_pooling(token_embeddings, mask):
    token_embeddings = token_embeddings.masked_fill(~mask[..., None].bool(), 0.)
    sentence_embeddings = token_embeddings.sum(dim=1) / mask.sum(dim=1)[..., None]
    return sentence_embeddings

embeddings = mean_pooling(outputs[0], sentences_inputs['attention_mask'])

需求是:通过向量点积计算句子相似度,为测试集中每个句子找出top_k(比如100个)最相似的ID,最终输出一个长度等于测试集的列表,每个元素是包含top_k个ID的子列表。

核心问题:内存限制导致无法同时存储超过约200个嵌入向量,如何完成上述计算?


解决方案

1. 分批次处理测试集嵌入

不要一次性生成所有测试集的向量,每次只处理一小批测试句子(比如200条,刚好匹配内存限制),生成对应的嵌入向量后,再用这批向量去遍历所有训练分块计算相似度,最后合并结果。内存中始终只保留一批测试向量,避免过载。

2. 逐块处理训练集,维护全局Top-K

对每一批测试向量,依次加载每个训练文件分块:

  • 加载当前训练分块的句子,同样分小批次生成嵌入向量(避免单个分块的向量占满内存)
  • 计算当前测试批次向量与训练分块小批次向量的点积矩阵(形状:[测试批次大小, 训练小批次大小])
  • 对每个测试样本,从当前训练小批次中取出top_k个最相似的ID和对应分数
  • 将这些局部top_k结果与之前的全局top_k结果合并,重新筛选出分数最高的100个ID

3. 内存优化细节

  • 训练分块内的小批次拆分:若单个训练分块的10万条句子生成的向量仍超内存,就拆成每批200条处理,计算完相似度后立即丢弃该批次的训练向量,再处理下一批。
  • 只保留必要数据:无需存储所有训练向量,计算完相似度后仅保留当前分块中每个测试样本的top_k ID和分数,大幅节省内存。
  • 关闭梯度计算:模型推理时使用torch.no_grad(),避免存储梯度信息,减少内存占用。
  • 及时清理缓存:处理完每个分块或批次后,手动删除变量并调用torch.cuda.empty_cache()释放GPU内存。

完整实现代码

import torch
from datasets import load_dataset
from transformers import AutoTokenizer, AutoModel
import heapq

# 初始化模型与分词器
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
tokenizer = AutoTokenizer.from_pretrained('facebook/contriever')
model = AutoModel.from_pretrained('facebook/contriever').to(device)
model.eval()  # 切换到评估模式

def get_embeddings(sentences, batch_size=200):
    """分批次生成句子嵌入,控制内存占用"""
    embeddings_list = []
    with torch.no_grad():  # 禁用梯度计算,节省内存
        for i in range(0, len(sentences), batch_size):
            batch_sentences = sentences[i:i+batch_size]
            inputs = tokenizer(
                batch_sentences,
                padding=True,
                truncation=True,
                return_tensors='pt'
            ).to(device)
            outputs = model(**inputs)
            batch_embeds = mean_pooling(outputs[0], inputs['attention_mask'])
            embeddings_list.append(batch_embeds.cpu())  # 移至CPU释放GPU空间
    return torch.cat(embeddings_list, dim=0)

def mean_pooling(token_embeddings, mask):
    token_embeddings = token_embeddings.masked_fill(~mask[..., None].bool(), 0.)
    sentence_embeddings = token_embeddings.sum(dim=1) / mask.sum(dim=1)[..., None]
    return sentence_embeddings

def update_top_k(current_top_k, new_scores, new_ids, top_k=100):
    """合并局部top_k结果到全局top_k,保留分数最高的项"""
    updated_top_k = []
    for existing, scores, ids in zip(current_top_k, new_scores, new_ids):
        # 合并已有结果与新结果
        combined = list(zip(scores.tolist(), ids.tolist())) + existing
        # 用堆排序快速筛选top_k
        top_items = heapq.nlargest(top_k, combined, key=lambda x: x[0])
        updated_top_k.append([item[1] for item in top_items])
    return updated_top_k

# 加载测试集
test_dataset = load_dataset('parquet', data_files='test_file.parquet')['train']
test_sentences = test_dataset['sentences']
total_test = len(test_sentences)
top_k = 100
test_batch_size = 200  # 匹配内存限制的批次大小

# 初始化最终结果列表
final_top_ids = [[(-float('inf'), -1)] * top_k for _ in range(total_test)]

# 分批次处理测试集
for test_start in range(0, total_test, test_batch_size):
    test_end = min(test_start + test_batch_size, total_test)
    test_batch = test_sentences[test_start:test_end]
    test_embeds = get_embeddings(test_batch).to(device)
    
    # 初始化当前测试批次的top_k结果
    current_batch_top = [[(-float('inf'), -1)] * top_k for _ in range(len(test_batch))]
    
    # 逐块处理所有训练分块
    for train_idx in range(30):
        # 加载单个训练分块
        train_chunk = load_dataset('parquet', data_files=f'file_{train_idx}.parquet')['train']
        train_sentences = train_chunk['sentences']
        train_ids = train_chunk['ids']
        
        # 分小批次处理训练分块
        train_batch_size = 200
        for train_start in range(0, len(train_sentences), train_batch_size):
            train_end = min(train_start + train_batch_size, len(train_sentences))
            train_batch_sentences = train_sentences[train_start:train_end]
            train_batch_ids = train_ids[train_start:train_end]
            
            # 获取训练批次的嵌入向量
            train_embeds = get_embeddings(train_batch_sentences).to(device)
            
            # 计算点积相似度矩阵
            scores = torch.matmul(test_embeds, train_embeds.T)  # shape: [test_batch_size, train_batch_size]
            
            # 获取当前训练批次的top_k分数与索引
            top_scores, top_indices = torch.topk(
                scores,
                k=min(top_k, len(train_batch_sentences)),
                dim=1
            )
            # 转换为对应的ID
            top_ids = [[train_batch_ids[i] for i in indices] for indices in top_indices.cpu().tolist()]
            
            # 更新当前测试批次的top_k结果
            current_batch_top = update_top_k(current_batch_top, top_scores, top_ids, top_k)
            
            # 释放训练批次的内存
            del train_embeds, scores, top_scores, top_indices, top_ids
            torch.cuda.empty_cache()
        
        # 释放训练分块的内存
        del train_chunk, train_sentences, train_ids
        torch.cuda.empty_cache()
    
    # 将当前测试批次的结果写入最终列表
    for i in range(len(current_batch_top)):
        final_top_ids[test_start + i] = current_batch_top[i]
    
    # 释放测试批次的内存
    del test_batch, test_embeds, current_batch_top
    torch.cuda.empty_cache()

# 输出结果示例(前5个测试样本的top_k ID)
print(final_top_ids[:5])

额外优化建议

  • 预存储训练集嵌入:若后续有重复计算需求,可预计算每个训练分块的嵌入并存储为.npy或.parquet文件,后续直接加载嵌入文件,避免重复运行模型编码,节省时间。Contriever向量为768维,10万条约307MB,30个文件共约9GB,需确认磁盘空间充足。
  • 使用近似最近邻算法:若对精度要求不严格,可采用FAISS、Annoy等近似最近邻库,大幅提升检索速度并降低内存占用。例如FAISS的IVF索引支持分块构建,适合大规模数据场景。

内容的提问来源于stack exchange,提问作者Penguin

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.11 07:12:05