Pinecone Upsert速度过慢求助:204条数据耗时2.5小时
问题诊断与提速方案
核心问题分析
- 模型未部署到GPU:代码仅定义了
device变量,但未将模型移动到GPU执行,嵌入计算全程在CPU运行,这是耗时过长的主要原因。 - 单条数据Upsert:每次仅向Pinecone Upsert一条数据,频繁的网络请求带来巨大开销。
- 冗余循环:代码先遍历
all_reviews[:2]做测试,之后又遍历全部数据,重复执行嵌入生成和Upsert操作。 - 单条文本嵌入生成:逐个处理文本块,未利用GPU并行计算能力,浪费硬件资源。
具体优化措施
- 将模型迁移到GPU:加载模型后执行
model = model.to(device),确保嵌入计算在GPU上运行。 - 批量生成嵌入:将多个文本块一次性传入模型生成嵌入,充分利用GPU并行性。
- 批量Upsert到Pinecone:积累一定数量的(ID、嵌入、元数据)三元组后,一次性调用
index.upsert(),减少网络请求次数。 - 移除冗余循环:删除测试用的
all_reviews[:2]遍历代码。
修改后的完整代码
初始化数据库
if index_name not in pc.list_indexes().names(): pc.create_index( name=index_name, dimension=1024, metric="cosine", spec=ServerlessSpec( cloud="aws", region="us-east-1" ) )
嵌入与批量Upsert操作
import torch from transformers import AutoTokenizer, AutoModel from langchain.text_splitter import SpacyTextSplitter device = 'cuda' if torch.cuda.is_available() else 'cpu' # 加载Jina嵌入模型与分词器 model_name = "jinaai/jina-embeddings-v3" tokenizer = AutoTokenizer.from_pretrained(model_name) model = AutoModel.from_pretrained(model_name, trust_remote_code=True) # 将模型移至GPU model = model.to(device) text_splitter = SpacyTextSplitter(chunk_size=500) # 批量生成嵌入函数 def generate_batch_embeddings(texts, task='retrieval.passage'): return model.encode(texts, convert_to_tensor=True, task=task, device=device).cpu().numpy() # 批量Upsert的批次大小,可根据GPU内存调整 batch_size = 64 upsert_batch = [] for review_id, review in enumerate(all_reviews): chunks = text_splitter.split_text(review) # 批量生成当前评论所有chunk的嵌入 embeddings = generate_batch_embeddings(chunks) for chunk_index, (chunk, embedding) in enumerate(zip(chunks, embeddings)): unique_id = f"{review_id}_{chunk_index}" metadata = {"review_id": review_id, "chunk_index": chunk_index, "text": chunk} upsert_batch.append((unique_id, embedding, metadata)) # 达到批次大小则执行Upsert if len(upsert_batch) >= batch_size: index.upsert(upsert_batch) upsert_batch = [] # 处理剩余的不足一个批次的数据 if upsert_batch: index.upsert(upsert_batch)
额外提速建议
- 调整批次大小:如果GPU内存充足,可适当增大
batch_size(如128或256),进一步提升嵌入生成效率。 - 替换文本分割工具:SpacyTextSplitter相对较慢,可替换为
RecursiveCharacterTextSplitter,减少文本分割耗时。 - 确认Colab GPU配置:确保已在Colab中启用GPU(菜单栏「Runtime」→「Change runtime type」→选择「GPU」)。
内容的提问来源于stack exchange,提问作者Daaku-C5
相关产品推荐
相关产品推荐

