You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

向Pinecone索引Upsert向量时遇numpy.ndarray类型错误求助

问题:Pinecone Upsert操作报错Invalid vector value passed: cannot interpret type <class 'numpy.ndarray'>

首次接触向量技术,参考网上代码将列表数据通过BERT生成向量嵌入后,向Pinecone索引执行upsert操作,代码如下:

tokenizer = BertTokenizer.from_pretrained("bert-base-uncased")
model = BertModel.from_pretrained("bert-base-uncased")

def generate_embeddings(text):
    inputs = tokenizer(text, return_tensors="pt")
    outputs = model(**inputs)
    embeddings = outputs.last_hidden_state.mean(dim=1).squeeze().detach().numpy()
    return embeddings

embeddings = [generate_embeddings(article) for article in article_content]


pc = Pinecone(api_key="API_KEY")
pc.create_index(
    name="index-name",
    dimension=4096, # Replace with your model dimensions
    metric="cosine", # Replace with your model metric
    spec=ServerlessSpec(
        cloud="aws",
        region="us-east-1"
    ) 
)
index = pc.Index("index-name")

index.upsert(embeddings)

运行时出现如下错误:

Traceback (most recent call last):
  File "c:\Users\tvish\llama api.py", line 158, in <module>
    index.upsert(embeddings)
  File "C:\Users\tvish\AppData\Local\Packages\PythonSoftwareFoundation.Python.3.11_qbz5n2kfra8p0\LocalCache\local-packages\Python311\site-packages\pinecone\utils\error_handling.py", line 11, in inner_func
    return func(*args, **kwargs)
           ^^^^^^^^^^^^^^^^^^^^^
  File "C:\Users\tvish\AppData\Local\Packages\PythonSoftwareFoundation.Python.3.11_qbz5n2kfra8p0\LocalCache\local-packages\Python311\site-packages\pinecone\data\index.py", line 175, in upsert
    return self._upsert_batch(vectors, namespace, _check_type, **kwargs)
           ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "C:\Users\tvish\AppData\Local\Packages\PythonSoftwareFoundation.Python.3.11_qbz5n2kfra8p0\LocalCache\local-packages\Python311\site-packages\pinecone\data\index.py", line 206, in _upsert_batch
    vectors=list(map(vec_builder, vectors)),
            ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "C:\Users\tvish\AppData\Local\Packages\PythonSoftwareFoundation.Python.3.11_qbz5n2kfra8p0\LocalCache\local-packages\Python311\site-packages\pinecone\data\index.py", line 202, in <lambda>
    vec_builder = lambda v: VectorFactory.build(v, check_type=_check_type)
                            ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "C:\Users\tvish\AppData\Local\Packages\PythonSoftwareFoundation.Python.3.11_qbz5n2kfra8p0\LocalCache\local-packages\Python311\site-packages\pinecone\data\vector_factory.py", line 30, in build
    raise ValueError(f"Invalid vector value passed: cannot interpret type {type(item)}")
ValueError: Invalid vector value passed: cannot interpret type <class 'numpy.ndarray'>

错误原因

  • Pinecone的upsert方法不接受纯numpy数组作为输入,它要求每个向量必须是包含唯一ID和向量值列表的字典结构
  • 代码中直接传入了numpy数组列表,既没有转换为Python原生列表,也没有添加必要的ID字段,导致Pinecone无法解析数据格式
  • 额外问题:bert-base-uncased生成的向量维度是768,代码中设置的dimension=4096完全不匹配,后续也会触发维度错误

解决方法

需要做三处关键修改:

  1. 将numpy数组转换为Python原生列表(用.tolist()方法)
  2. 为每个向量生成唯一ID(可用索引序号或UUID)
  3. 修正索引维度为BERT模型对应的768

修改后的完整代码:

import uuid
from transformers import BertTokenizer, BertModel
from pinecone import Pinecone, ServerlessSpec

tokenizer = BertTokenizer.from_pretrained("bert-base-uncased")
model = BertModel.from_pretrained("bert-base-uncased")

def generate_embeddings(text):
    inputs = tokenizer(text, return_tensors="pt")
    outputs = model(**inputs)
    # 将numpy数组转为Python原生列表
    embeddings = outputs.last_hidden_state.mean(dim=1).squeeze().detach().numpy().tolist()
    return embeddings

# 构造符合Pinecone要求的向量字典列表
vectors = [
    {
        "id": str(uuid.uuid4()),  # 生成唯一UUID作为向量ID
        "values": generate_embeddings(article),
        # 可选:添加元数据,比如{"original_text": article}
    }
    for article in article_content
]

pc = Pinecone(api_key="API_KEY")
# 注意:如果索引已存在,删除该行避免重复创建报错
pc.create_index(
    name="index-name",
    dimension=768,  # 修正为bert-base-uncased的正确维度
    metric="cosine",
    spec=ServerlessSpec(
        cloud="aws",
        region="us-east-1"
    ) 
)
index = pc.Index("index-name")

# 传入正确格式的向量列表执行upsert
index.upsert(vectors=vectors)

额外提示:

  • 如果索引已经创建完成,务必删除pc.create_index()代码行,重复创建会触发索引已存在的错误
  • 若需要批量处理大量数据,可以分批次调用upsert,避免单次请求数据量过大触发限流

内容的提问来源于stack exchange,提问作者TvishaCat

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.19 04:25:15