向Pinecone索引Upsert向量时遇numpy.ndarray类型错误求助
问题:Pinecone Upsert操作报错
Invalid vector value passed: cannot interpret type <class 'numpy.ndarray'> 首次接触向量技术,参考网上代码将列表数据通过BERT生成向量嵌入后,向Pinecone索引执行upsert操作,代码如下:
tokenizer = BertTokenizer.from_pretrained("bert-base-uncased") model = BertModel.from_pretrained("bert-base-uncased") def generate_embeddings(text): inputs = tokenizer(text, return_tensors="pt") outputs = model(**inputs) embeddings = outputs.last_hidden_state.mean(dim=1).squeeze().detach().numpy() return embeddings embeddings = [generate_embeddings(article) for article in article_content] pc = Pinecone(api_key="API_KEY") pc.create_index( name="index-name", dimension=4096, # Replace with your model dimensions metric="cosine", # Replace with your model metric spec=ServerlessSpec( cloud="aws", region="us-east-1" ) ) index = pc.Index("index-name") index.upsert(embeddings)
运行时出现如下错误:
Traceback (most recent call last): File "c:\Users\tvish\llama api.py", line 158, in <module> index.upsert(embeddings) File "C:\Users\tvish\AppData\Local\Packages\PythonSoftwareFoundation.Python.3.11_qbz5n2kfra8p0\LocalCache\local-packages\Python311\site-packages\pinecone\utils\error_handling.py", line 11, in inner_func return func(*args, **kwargs) ^^^^^^^^^^^^^^^^^^^^^ File "C:\Users\tvish\AppData\Local\Packages\PythonSoftwareFoundation.Python.3.11_qbz5n2kfra8p0\LocalCache\local-packages\Python311\site-packages\pinecone\data\index.py", line 175, in upsert return self._upsert_batch(vectors, namespace, _check_type, **kwargs) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "C:\Users\tvish\AppData\Local\Packages\PythonSoftwareFoundation.Python.3.11_qbz5n2kfra8p0\LocalCache\local-packages\Python311\site-packages\pinecone\data\index.py", line 206, in _upsert_batch vectors=list(map(vec_builder, vectors)), ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "C:\Users\tvish\AppData\Local\Packages\PythonSoftwareFoundation.Python.3.11_qbz5n2kfra8p0\LocalCache\local-packages\Python311\site-packages\pinecone\data\index.py", line 202, in <lambda> vec_builder = lambda v: VectorFactory.build(v, check_type=_check_type) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "C:\Users\tvish\AppData\Local\Packages\PythonSoftwareFoundation.Python.3.11_qbz5n2kfra8p0\LocalCache\local-packages\Python311\site-packages\pinecone\data\vector_factory.py", line 30, in build raise ValueError(f"Invalid vector value passed: cannot interpret type {type(item)}") ValueError: Invalid vector value passed: cannot interpret type <class 'numpy.ndarray'>
错误原因
- Pinecone的
upsert方法不接受纯numpy数组作为输入,它要求每个向量必须是包含唯一ID和向量值列表的字典结构 - 代码中直接传入了numpy数组列表,既没有转换为Python原生列表,也没有添加必要的ID字段,导致Pinecone无法解析数据格式
- 额外问题:
bert-base-uncased生成的向量维度是768,代码中设置的dimension=4096完全不匹配,后续也会触发维度错误
解决方法
需要做三处关键修改:
- 将numpy数组转换为Python原生列表(用
.tolist()方法) - 为每个向量生成唯一ID(可用索引序号或UUID)
- 修正索引维度为BERT模型对应的768
修改后的完整代码:
import uuid from transformers import BertTokenizer, BertModel from pinecone import Pinecone, ServerlessSpec tokenizer = BertTokenizer.from_pretrained("bert-base-uncased") model = BertModel.from_pretrained("bert-base-uncased") def generate_embeddings(text): inputs = tokenizer(text, return_tensors="pt") outputs = model(**inputs) # 将numpy数组转为Python原生列表 embeddings = outputs.last_hidden_state.mean(dim=1).squeeze().detach().numpy().tolist() return embeddings # 构造符合Pinecone要求的向量字典列表 vectors = [ { "id": str(uuid.uuid4()), # 生成唯一UUID作为向量ID "values": generate_embeddings(article), # 可选:添加元数据,比如{"original_text": article} } for article in article_content ] pc = Pinecone(api_key="API_KEY") # 注意:如果索引已存在,删除该行避免重复创建报错 pc.create_index( name="index-name", dimension=768, # 修正为bert-base-uncased的正确维度 metric="cosine", spec=ServerlessSpec( cloud="aws", region="us-east-1" ) ) index = pc.Index("index-name") # 传入正确格式的向量列表执行upsert index.upsert(vectors=vectors)
额外提示:
- 如果索引已经创建完成,务必删除
pc.create_index()代码行,重复创建会触发索引已存在的错误 - 若需要批量处理大量数据,可以分批次调用
upsert,避免单次请求数据量过大触发限流
内容的提问来源于stack exchange,提问作者TvishaCat
相关产品推荐
相关产品推荐

