You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pinecone Upsert速度过慢求助:204条数据耗时2.5小时

问题诊断与提速方案

核心问题分析

  1. 模型未部署到GPU:代码仅定义了device变量,但未将模型移动到GPU执行,嵌入计算全程在CPU运行,这是耗时过长的主要原因。
  2. 单条数据Upsert:每次仅向Pinecone Upsert一条数据,频繁的网络请求带来巨大开销。
  3. 冗余循环:代码先遍历all_reviews[:2]做测试,之后又遍历全部数据,重复执行嵌入生成和Upsert操作。
  4. 单条文本嵌入生成:逐个处理文本块,未利用GPU并行计算能力,浪费硬件资源。

具体优化措施

  • 将模型迁移到GPU:加载模型后执行model = model.to(device),确保嵌入计算在GPU上运行。
  • 批量生成嵌入:将多个文本块一次性传入模型生成嵌入,充分利用GPU并行性。
  • 批量Upsert到Pinecone:积累一定数量的(ID、嵌入、元数据)三元组后,一次性调用index.upsert(),减少网络请求次数。
  • 移除冗余循环:删除测试用的all_reviews[:2]遍历代码。

修改后的完整代码

初始化数据库

if index_name not in pc.list_indexes().names():
  pc.create_index(
    name=index_name,
    dimension=1024,
    metric="cosine",
    spec=ServerlessSpec(
        cloud="aws",
        region="us-east-1"
    )
)

嵌入与批量Upsert操作

import torch
from transformers import AutoTokenizer, AutoModel
from langchain.text_splitter import SpacyTextSplitter

device = 'cuda' if torch.cuda.is_available() else 'cpu'

# 加载Jina嵌入模型与分词器
model_name = "jinaai/jina-embeddings-v3"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModel.from_pretrained(model_name, trust_remote_code=True)
# 将模型移至GPU
model = model.to(device)

text_splitter = SpacyTextSplitter(chunk_size=500)

# 批量生成嵌入函数
def generate_batch_embeddings(texts, task='retrieval.passage'):
    return model.encode(texts, convert_to_tensor=True, task=task, device=device).cpu().numpy()

# 批量Upsert的批次大小,可根据GPU内存调整
batch_size = 64
upsert_batch = []

for review_id, review in enumerate(all_reviews):
    chunks = text_splitter.split_text(review)
    
    # 批量生成当前评论所有chunk的嵌入
    embeddings = generate_batch_embeddings(chunks)
    
    for chunk_index, (chunk, embedding) in enumerate(zip(chunks, embeddings)):
        unique_id = f"{review_id}_{chunk_index}"
        metadata = {"review_id": review_id, "chunk_index": chunk_index, "text": chunk}
        upsert_batch.append((unique_id, embedding, metadata))
        
        # 达到批次大小则执行Upsert
        if len(upsert_batch) >= batch_size:
            index.upsert(upsert_batch)
            upsert_batch = []

# 处理剩余的不足一个批次的数据
if upsert_batch:
    index.upsert(upsert_batch)

额外提速建议

  • 调整批次大小:如果GPU内存充足,可适当增大batch_size(如128或256),进一步提升嵌入生成效率。
  • 替换文本分割工具:SpacyTextSplitter相对较慢,可替换为RecursiveCharacterTextSplitter,减少文本分割耗时。
  • 确认Colab GPU配置:确保已在Colab中启用GPU(菜单栏「Runtime」→「Change runtime type」→选择「GPU」)。

内容的提问来源于stack exchange,提问作者Daaku-C5

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.13 23:44:51