You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

异步调用Gemini-1.5-Flash处理文本时遭遇429资源耗尽错误的排查求助

异步调用Gemini-1.5-Flash处理文本时遭遇429资源耗尽错误的排查求助

看起来你在异步批量处理4个文本分片时,调用Gemini-1.5-Flash频繁触发429资源耗尽错误,虽然后台配额显示允许2000并发请求,但实际还是踩了限制。我来帮你梳理几个实用的排查和解决方向:

一、先搞清楚真正触发限制的原因

你看到的“2000并发请求”配额,往往不是唯一的限制条件。Gemini的配额通常分三类,429大概率是后两类触发的:

  • 并发请求数:同时运行的请求数量
  • 请求速率限制:每分钟/每秒允许的请求总数(QPS/RPM)
  • token总量限制:每分钟/每秒允许处理的token总数(输入+输出)

建议你去Google Cloud控制台的Gemini配额详情页,仔细核对这三类限制的当前使用情况——很多时候是短时间内请求集中发送,触发了QPS限制,而非并发数超标。

二、给异步请求加个“节流阀”

哪怕配额写了2000并发,云服务也不建议一次性把所有请求都砸出去。你可以用asyncio.Semaphore控制同时运行的任务数,比如先限制到5个以内试试:

# 在创建tasks前定义信号量
semaphore = asyncio.Semaphore(5)

# 修改process_chunk函数,加入信号量控制
async def process_chunk(idx: int, text: str) -> (int, str):
    async with semaphore:
        # 原有的prompt构建、请求逻辑...
        response = await client.chat.completions.create(
            model=model,
            temperature=0,
            response_format={"type": "text"},
            messages=[
                {"role": "system", "content": system_prompt},
                {"role": "user", "content": user_prompt}
            ]
        )
        # 后续处理逻辑...

如果还是触发限制,可以再降低并发数,或者在任务之间加个微小延迟(比如await asyncio.sleep(0.1)),避免请求集中爆发。

三、换成官方异步客户端更稳妥

你代码里用AsyncOpenAI兼容层调用Gemini,可能存在适配问题。建议换成Google官方的google-generativeai异步客户端,它能更好地处理Gemini的限流规则,甚至自带重试机制:

import google.generativeai as genai

# 初始化客户端
genai.configure(api_key=GEMINI_KEY)
gemini_model = genai.GenerativeModel('gemini-1.5-flash')

async def process_chunk(idx: int, text: str) -> (int, str):
    print(f"Chunk {idx} processing...")
    user_prompt = (
        f"Please rewrite the following text with proper capitalization, punctuation, and line breaks "
        f"without adding or removing any words or content:\n\n{text}"
    )
    # 异步调用Gemini
    response = await gemini_model.generate_content_async(
        [system_prompt, user_prompt],
        temperature=0
    )
    result = response.text
    print(f"Chunk {idx} processed")
    return idx, result

四、给请求加自动重试机制

429很多时候是临时限流,你可以用tenacity库给请求加自动重试逻辑,避免直接失败:

from tenacity import retry, stop_after_attempt, wait_exponential, retry_if_exception_type

@retry(
    stop=stop_after_attempt(3),  # 最多重试3次
    wait=wait_exponential(multiplier=1, min=2, max=10),  # 指数退避等待
    retry=retry_if_exception_type(Exception)  # 遇到异常就重试
)
async def process_chunk(idx: int, text: str) -> (int, str):
    # 原有的请求逻辑...

五、检查文本分片的token计算是否准确

Gemini的token计算规则和OpenAI不一样,你当前的split_text_by_token_limit函数如果用的是OpenAI的tiktoken计算,可能导致分片实际token数超标,触发单请求的资源限制。建议换成Gemini官方的计数方法:

# 用Gemini官方方法计算token数
token_count = genai.count_tokens(text)

你可以先从核查配额详情和添加请求节流这两步入手,这是最常见的解决方向。如果还是不行,再尝试替换官方客户端和添加重试机制。

备注:内容来源于stack exchange,提问作者Cheok Yan Cheng

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.14 10:54:31