You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

GCP免费期用Langchain4J调用Vertex API文本嵌入遇配额超限问题

解决GCP Vertex API文本嵌入配额超限问题

我正尝试在Micronaut应用中通过Langchain4J调用Vertex API实现文本嵌入功能,代码实现如下:

@Singleton
public record VertexAiEmbedding(GoogleCloudConfiguration googleCloudConfiguration, VertexAiConfig vertexAiConfig) implements IVertexAiEmbedding {
    private static Embedding embedding;
    @Override
    public float[] embedVector(String text) {
        EmbeddingModel embeddingModel = VertexAiEmbeddingModel.builder()
                .endpoint(vertexAiConfig.endPoint())
                .project(googleCloudConfiguration.getProjectId())
                .location(vertexAiConfig.location())
                .publisher(vertexAiConfig.publisher())
                .modelName(vertexAiConfig.modelName())
                .build();
        Response<Embedding> response = embeddingModel.embed(text);
        embedding = response.content();
        return embedding.vector();
    }
}

运行时触发异常:

Caused by: com.google.api.gax.rpc.ResourceExhaustedException: io.grpc.StatusRuntimeException: RESOURCE_EXHAUSTED: Quota exceeded for quota metric 'LLM utility requests' and limit 'LLM utility requests per minute per region' of service 'aiplatform.googleapis.com' for consumer 'project_number:974067563912'.

当前处于GCP免费试用阶段,相关额度截图:
配额截图

更新后的账户额度截图:
账户额度截图1
账户额度截图2


解决方案

1. 优化请求频率,添加限流

异常明确是每分钟请求数超过配额限制,免费试用阶段配额通常较低,可通过限流控制请求频次:

  • 使用Guava的RateLimiter在调用层限制请求速率,示例代码:
    import com.google.common.util.concurrent.RateLimiter;
    
    @Singleton
    public record VertexAiEmbedding(GoogleCloudConfiguration googleCloudConfiguration, VertexAiConfig vertexAiConfig) implements IVertexAiEmbedding {
        private static Embedding embedding;
        private final RateLimiter rateLimiter = RateLimiter.create(1.0); // 每秒1次请求,对应每分钟60次,可根据配额调整
        private final EmbeddingModel embeddingModel;
    
        public VertexAiEmbedding {
            embeddingModel = VertexAiEmbeddingModel.builder()
                    .endpoint(vertexAiConfig.endPoint())
                    .project(googleCloudConfiguration.getProjectId())
                    .location(vertexAiConfig.location())
                    .publisher(vertexAiConfig.publisher())
                    .modelName(vertexAiConfig.modelName())
                    .build();
        }
    
        @Override
        public float[] embedVector(String text) {
            rateLimiter.acquire(); // 阻塞等待直到获取请求许可
            Response<Embedding> response = embeddingModel.embed(text);
            embedding = response.content();
            return embedding.vector();
        }
    }
    
  • 若存在批量文本处理场景,优先使用Langchain4J的embedAll批量嵌入方法,减少单次请求数量,降低请求频次。

2. 申请提升配额

免费试用账户可申请提升配额以满足测试需求:

  • 进入GCP控制台,搜索「配额」进入配额管理页面;
  • 找到aiplatform.googleapis.com服务下的「LLM utility requests per minute per region」配额项;
  • 点击「编辑配额」,填写合理的申请理由(如开发测试需要)并提交,通常几小时到1天内会完成审核。

3. 优化代码冗余

当前代码每次调用都重新构建EmbeddingModel实例,既浪费资源也可能间接增加请求开销,改为构造时初始化单例实例(如上优化后的代码所示)。


内容的提问来源于stack exchange,提问作者San Jaisy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.26 01:15:26