You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

GCP Vertex AI RAG语料库文件上传缓慢及地域匹配报错问题求助

GCP Vertex AI RAG语料库文件上传缓慢及地域匹配报错问题求助

最近在GCP Vertex AI上做RAG相关的实验,写了个简单的脚本:先本地生成几个带Lorem Ipsum内容的测试文件,再逐个上传到新建的RAG语料库。脚本代码如下:

import vertexai
from vertexai import rag
from tqdm import tqdm
from pathlib import Path
import lorem

PROJECT_ID = "my-project-id"  # 按需修改
LOCATION = "us-central1"
CORPUS_DISPLAY_NAME = f"dummy_corpus"
TEMP_FILES_DIR = Path("temp_rag_files")


def create_files(num_files=5):
    """Creates a specified number of dummy text files with lorem ipsum content."""
    TEMP_FILES_DIR.mkdir(exist_ok=True)
    created_file_paths = []
    for i in range(num_files):
        file_path = TEMP_FILES_DIR / f"dummy_file_{i+1}.txt"
        content = f"Dummy file {i+1} for RAG example.\n{lorem.paragraph()}"
        file_path.write_text(content, encoding='utf-8')
        created_file_paths.append(file_path)
        print(f"Created dummy file: {file_path}")
    return created_file_paths


def main():
    vertexai.init(project=PROJECT_ID, location=LOCATION)
    print("Creating dummy files...")
    dummy_file_paths = create_files(num_files=5)

    print(f"Creating RAG Corpus '{CORPUS_DISPLAY_NAME}'...")
    corpus = rag.create_corpus(
        display_name=CORPUS_DISPLAY_NAME,
        description="Corpus with lorem ipsum files.",
    )
    corpus_name = corpus.name
    print(f"Successfully created RAG Corpus: {corpus_name}")

    print(f"Uploading {len(dummy_file_paths)} files to '{corpus_name}'...")
    uploaded_rag_files_info = []
    for file_path in tqdm(dummy_file_paths):
        display_name = file_path.stem
        rag_file = rag.upload_file(
            corpus_name=corpus_name,
            path=str(file_path),
            display_name=display_name,
            description=f"Dummy lorem ipsum file: {display_name}",
        )
        uploaded_rag_files_info.append({"name": rag_file.name, "display_name": rag_file.display_name})
        print(f"Successfully uploaded: {rag_file.name}")


if __name__ == "__main__":
    main()

不过运行的时候遇到了两个棘手的问题,想请教下大家:

  1. 文件上传速度极慢:每个文件的上传流程要花大概10秒,要是数据集稍微大一点,这种上传方式完全不可行。我是不是哪里写法不对?有没有更高效的批量上传方式?

  2. 地域匹配报错:一开始以为慢是因为地域延迟,我把LOCATION改成了离我更近的europe-west3,结果直接报错:

RuntimeError: ('Failed in indexing the RagFile due to: ', {'code': 400, 'message': 'Request resource location europe-west3 does not match service location us-central1.', 'status': 'FAILED_PRECONDITION'})

这个报错我有点懵,难道RAG服务目前只支持us-central1地域?还是我初始化Vertex AI的时候有什么配置遗漏了?

目前我在考虑换其他上传方式,但还是想先搞清楚这两个问题的原因,麻烦各位大佬指点下!

内容来源于stack exchange

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.08 11:09:35