使用ChatGPT检索插件向Pinecone上传元数据失败问题排查
问题:ChatGPT检索插件上传文件时元数据字段为空的解决方法
我使用ChatGPT检索插件结合Pinecone向量数据库实现自定义数据的向量化存储,文件可正常上传并入库,但自定义的元数据(如来源、作者、URL等)大部分字段返回None。
当前上传代码
def upsert_file(directory: str): """ Upload all files under a directory to the vector database. """ url = "http://0.0.0.0:8000/upsert-file" headers = {"Authorization": "Bearer " + DATABASE_INTERFACE_BEARER_TOKEN} files = [] for filename in os.listdir(directory): if os.path.isfile(os.path.join(directory, filename)): file_path = os.path.join(directory, filename) with open(file_path, "rb") as f: file_content = f.read() metadata = { "source": filename, "author": "Tim Cook", "url": "Some fake url" } print(metadata) files = { "file": (filename, file_content, "text/plain"), "metadata": (None, json.dumps(metadata), "application/json"), } response = requests.post(url, headers=headers, files=files, timeout=600) if response.status_code == 200: print(filename + " uploaded successfully.") else: print( f"Error: {response.status_code} {response.content} for uploading " + filename)
遇到的问题
- 元数据丢失:文件正常存储后,返回的元数据多数字段为
None:
metadata': {'source': 'file', 'source_id': None, 'url': None, 'created_at': None, 'author': None, 'document_id': 'Some_Doc_Id_here_that_is_not_None'}
- 参数类型错误:若将
"metadata": (None, json.dumps(metadata), "application/json")中的None改为字符串(如testing),触发422错误:
Error: 422 b'{"detail":[{"loc":["body","metadata"],"msg":"str type expected","type":"type_error.str"}]}'
解决方法
1. 调整参数传递方式
ChatGPT检索插件的upsert-file接口要求元数据通过data参数传递,而非作为文件类型放入files字典。接口期望metadata是JSON格式的字符串,不是文件对象。
2. 修改后的正确代码
import os import json import requests def upsert_file(directory: str): """ Upload all files under a directory to the vector database. """ url = "http://0.0.0.0:8000/upsert-file" headers = {"Authorization": "Bearer " + DATABASE_INTERFACE_BEARER_TOKEN} for filename in os.listdir(directory): file_path = os.path.join(directory, filename) if not os.path.isfile(file_path): continue with open(file_path, "rb") as f: file_content = f.read() metadata = { "source": filename, "author": "Tim Cook", "url": "Some fake url" } print(metadata) # 文件放入files,元数据转JSON字符串放入data files = { "file": (filename, file_content, "text/plain") } data = { "metadata": json.dumps(metadata) } response = requests.post( url, headers=headers, files=files, data=data, timeout=600 ) if response.status_code == 200: print(f"{filename} uploaded successfully.") else: print(f"Error: {response.status_code} {response.content} for uploading {filename}")
3. 问题原因说明
- 之前将
metadata放在files中,插件会将其识别为文件,无法解析为元数据对象,导致字段丢失; - 修改
None为字符串后,插件收到的是文件类型参数,但接口要求metadata是字符串类型的JSON,因此触发类型错误。
4. 验证
修改代码后重新上传文件,查询Pinecone中的数据,自定义元数据字段会被正确填充。
内容的提问来源于stack exchange,提问作者Mark
相关产品推荐
相关产品推荐

