You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用ChatGPT检索插件向Pinecone上传元数据失败问题排查

问题:ChatGPT检索插件上传文件时元数据字段为空的解决方法

我使用ChatGPT检索插件结合Pinecone向量数据库实现自定义数据的向量化存储,文件可正常上传并入库,但自定义的元数据(如来源、作者、URL等)大部分字段返回None。

当前上传代码

def upsert_file(directory: str):
    """
    Upload all files under a directory to the vector database.
    """
    url = "http://0.0.0.0:8000/upsert-file"
    headers = {"Authorization": "Bearer " + DATABASE_INTERFACE_BEARER_TOKEN}
    files = []
    for filename in os.listdir(directory):
        if os.path.isfile(os.path.join(directory, filename)):
            file_path = os.path.join(directory, filename)
            with open(file_path, "rb") as f:
                file_content = f.read()
                metadata = {
                    "source": filename,
                    "author": "Tim Cook",
                    "url": "Some fake url"
                }
                print(metadata)
                files = {
                    "file": (filename, file_content, "text/plain"),
                    "metadata": (None, json.dumps(metadata), "application/json"),
                }

            response = requests.post(url,
                                     headers=headers,
                                     files=files,
                                     timeout=600)
            if response.status_code == 200:
                print(filename + " uploaded successfully.")
            else:
                print(
                    f"Error: {response.status_code} {response.content} for uploading "
                    + filename)

遇到的问题

  1. 元数据丢失:文件正常存储后,返回的元数据多数字段为None:
metadata': {'source': 'file', 'source_id': None, 'url': None, 'created_at': None, 'author': None, 'document_id': 'Some_Doc_Id_here_that_is_not_None'}
  1. 参数类型错误:若将"metadata": (None, json.dumps(metadata), "application/json")中的None改为字符串(如testing),触发422错误:
Error: 422 b'{"detail":[{"loc":["body","metadata"],"msg":"str type expected","type":"type_error.str"}]}'

解决方法

1. 调整参数传递方式

ChatGPT检索插件的upsert-file接口要求元数据通过data参数传递,而非作为文件类型放入files字典。接口期望metadata是JSON格式的字符串,不是文件对象。

2. 修改后的正确代码

import os
import json
import requests

def upsert_file(directory: str):
    """
    Upload all files under a directory to the vector database.
    """
    url = "http://0.0.0.0:8000/upsert-file"
    headers = {"Authorization": "Bearer " + DATABASE_INTERFACE_BEARER_TOKEN}
    
    for filename in os.listdir(directory):
        file_path = os.path.join(directory, filename)
        if not os.path.isfile(file_path):
            continue
            
        with open(file_path, "rb") as f:
            file_content = f.read()
            metadata = {
                "source": filename,
                "author": "Tim Cook",
                "url": "Some fake url"
            }
            print(metadata)
            # 文件放入files,元数据转JSON字符串放入data
            files = {
                "file": (filename, file_content, "text/plain")
            }
            data = {
                "metadata": json.dumps(metadata)
            }

            response = requests.post(
                url,
                headers=headers,
                files=files,
                data=data,
                timeout=600
            )
            
            if response.status_code == 200:
                print(f"{filename} uploaded successfully.")
            else:
                print(f"Error: {response.status_code} {response.content} for uploading {filename}")

3. 问题原因说明

  • 之前将metadata放在files中,插件会将其识别为文件,无法解析为元数据对象,导致字段丢失;
  • 修改None为字符串后,插件收到的是文件类型参数,但接口要求metadata是字符串类型的JSON,因此触发类型错误。

4. 验证

修改代码后重新上传文件,查询Pinecone中的数据,自定义元数据字段会被正确填充。


内容的提问来源于stack exchange,提问作者Mark

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.16 21:55:12