GCP免费期用Langchain4J调用Vertex API文本嵌入遇配额超限问题
解决GCP Vertex API文本嵌入配额超限问题
我正尝试在Micronaut应用中通过Langchain4J调用Vertex API实现文本嵌入功能,代码实现如下:
@Singleton public record VertexAiEmbedding(GoogleCloudConfiguration googleCloudConfiguration, VertexAiConfig vertexAiConfig) implements IVertexAiEmbedding { private static Embedding embedding; @Override public float[] embedVector(String text) { EmbeddingModel embeddingModel = VertexAiEmbeddingModel.builder() .endpoint(vertexAiConfig.endPoint()) .project(googleCloudConfiguration.getProjectId()) .location(vertexAiConfig.location()) .publisher(vertexAiConfig.publisher()) .modelName(vertexAiConfig.modelName()) .build(); Response<Embedding> response = embeddingModel.embed(text); embedding = response.content(); return embedding.vector(); } }
运行时触发异常:
Caused by: com.google.api.gax.rpc.ResourceExhaustedException: io.grpc.StatusRuntimeException: RESOURCE_EXHAUSTED: Quota exceeded for quota metric 'LLM utility requests' and limit 'LLM utility requests per minute per region' of service 'aiplatform.googleapis.com' for consumer 'project_number:974067563912'.
当前处于GCP免费试用阶段,相关额度截图:
更新后的账户额度截图:

解决方案
1. 优化请求频率,添加限流
异常明确是每分钟请求数超过配额限制,免费试用阶段配额通常较低,可通过限流控制请求频次:
- 使用Guava的
RateLimiter在调用层限制请求速率,示例代码:import com.google.common.util.concurrent.RateLimiter; @Singleton public record VertexAiEmbedding(GoogleCloudConfiguration googleCloudConfiguration, VertexAiConfig vertexAiConfig) implements IVertexAiEmbedding { private static Embedding embedding; private final RateLimiter rateLimiter = RateLimiter.create(1.0); // 每秒1次请求,对应每分钟60次,可根据配额调整 private final EmbeddingModel embeddingModel; public VertexAiEmbedding { embeddingModel = VertexAiEmbeddingModel.builder() .endpoint(vertexAiConfig.endPoint()) .project(googleCloudConfiguration.getProjectId()) .location(vertexAiConfig.location()) .publisher(vertexAiConfig.publisher()) .modelName(vertexAiConfig.modelName()) .build(); } @Override public float[] embedVector(String text) { rateLimiter.acquire(); // 阻塞等待直到获取请求许可 Response<Embedding> response = embeddingModel.embed(text); embedding = response.content(); return embedding.vector(); } } - 若存在批量文本处理场景,优先使用Langchain4J的
embedAll批量嵌入方法,减少单次请求数量,降低请求频次。
2. 申请提升配额
免费试用账户可申请提升配额以满足测试需求:
- 进入GCP控制台,搜索「配额」进入配额管理页面;
- 找到
aiplatform.googleapis.com服务下的「LLM utility requests per minute per region」配额项; - 点击「编辑配额」,填写合理的申请理由(如开发测试需要)并提交,通常几小时到1天内会完成审核。
3. 优化代码冗余
当前代码每次调用都重新构建EmbeddingModel实例,既浪费资源也可能间接增加请求开销,改为构造时初始化单例实例(如上优化后的代码所示)。
内容的提问来源于stack exchange,提问作者San Jaisy
相关产品推荐
相关产品推荐

