You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Cohere API的Streamlit应用遭遇Token超限错误求助

解决CohereAPIError: too many tokens 问题

问题说明

基于Streamlit开发的自然语言处理应用调用Cohere API时触发错误:CohereAPIError:too many tokens: total number of tokens in the prompt cannot exceed 4081 - received 15416,核心原因是请求的prompt令牌总数超出了API设定的4081令牌上限。

可行解决方案

  • 截断超长Prompt:利用Cohere的令牌化工具计算prompt的令牌数量,超出上限时直接截断文本。示例代码:

    import cohere
    co = cohere.Client("你的API密钥")
    
    def truncate_prompt(prompt, max_tokens=4081):
        tokenized_data = co.tokenize(text=prompt)
        if len(tokenized_data.tokens) > max_tokens:
            truncated_tokens = tokenized_data.tokens[:max_tokens]
            return co.detokenize(truncated_tokens).text
        return prompt
    
    # 使用示例
    user_raw_prompt = "用户输入的超长内容..."
    safe_prompt = truncate_prompt(user_raw_prompt)
    
  • 精简Prompt内容:移除prompt中冗余的背景描述、重复指令等非必要信息,仅保留核心任务要求和用户问题。比如处理对话类请求时,只保留最近几轮关键交互内容。

  • 分批次处理长文本:若需处理超长文档,将文本分割为多个不超过令牌上限的块,分别调用API后合并结果。示例思路:

    def split_text_into_chunks(text, max_tokens=4081):
        tokens = co.tokenize(text=text).tokens
        chunks = []
        for i in range(0, len(tokens), max_tokens):
            chunk_tokens = tokens[i:i+max_tokens]
            chunks.append(co.detokenize(chunk_tokens).text)
        return chunks
    
    # 分块处理并合并结果
    long_target_text = "需要处理的超长文本内容..."
    text_chunks = split_text_into_chunks(long_target_text)
    processed_results = []
    for chunk in text_chunks:
        api_response = co.generate(prompt=chunk)
        processed_results.append(api_response.generations[0].text)
    final_result = "\n".join(processed_results)
    
  • 切换大上下文模型:Cohere部分模型(如Command R)支持更高的令牌上限,若业务场景允许,可切换至这类模型规避超限问题。

内容的提问来源于stack exchange,提问作者Davd Anderson

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.04 13:21:07