使用OpenAI Python库上传JSON文件至助手向量存储时的文件处理异常排查请求
OpenAI Python库上传JSON文件至助手向量存储时的文件处理异常排查请求
我现在尝试用OpenAI官方Python库把两个JSON文件上传到助手的向量存储里,还想给每个文件设置不同的分块策略。
我找到了两种可行的上传文件到向量存储的方式,但两种都出问题了:
- 第一种是先创建向量存储,再用
client.beta.vector_stores.files相关函数上传文件。代码执行完全没抛出异常,但向量存储是空的,显示0个文件。 - 第二种是先用
client.files.create上传文件,再创建向量存储并关联之前上传的文件。同样没异常,但向量存储的file_count一直停留在in_progress = 2的状态,卡住不动了。
第一种方法的代码(向量存储最终为空)
我用这段代码创建向量存储,然后给每个文件指定分块策略上传,结果向量存储是空的,没有任何异常:
vector_store = client.beta.vector_stores.create( name="human labeled dataset", ) client.beta.vector_stores.files.upload_and_poll( vector_store_id=vector_store.id, file=open("results/results_tsm_human_labeled.json", "rb"), poll_interval_ms=1000, chunking_strategy={ "type": "static", "static": {"max_chunk_size_tokens": 100, "chunk_overlap_tokens": 5}, }, ) client.beta.vector_stores.files.upload_and_poll( vector_store_id=vector_store.id, file=open("data/sample_tsm_new.json", "rb"), poll_interval_ms=1000, chunking_strategy={ "type": "static", "static": {"max_chunk_size_tokens": 1000, "chunk_overlap_tokens": 400}, }, )
就算去掉分块策略的部分,结果还是一样,向量存储依旧是空的。
第二种方法的代码(状态卡住)
为了调试,我试了先上传文件再创建向量存储的方式,没加任何分块策略:
human_dataset_result_json_file = client.files.create( file=open("results/results_tsm_human_labeled.json", "rb"), purpose="assistants" ) human_dataset_json_file = client.files.create( file=open("data/sample_tsm_new.json", "rb"), purpose="assistants" ) vector_store = client.beta.vector_stores.create( name="human labeled dataset", file_ids=[human_dataset_result_json_file.id, human_dataset_json_file.id], )
结果这个向量存储的文件处理状态一直卡在in_progress = 2,再也没变化。
有意思的是,通过Web UI把这两个文件上传到向量存储完全正常,没有任何问题。
为什么会出现这种情况呢?
备注:内容来源于stack exchange,提问作者user28146142
相关产品推荐
相关产品推荐

