You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用DocumentSummaryIndex时Llama Index TokenCountingHandler始终返回0值

Llama Index中DocumentSummaryIndex的TokenCountingHandler计数异常问题

使用Llama Index时遇到token计数异常:配置了MockLLM、MockEmbedding及TokenCountingHandler并加入CallbackManager,通过get_response_synthesizer设置tree_summarize模式响应合成器,创建DocumentSummaryIndex后,打印token计数器的各项数值均为0。疑问Llama Index的token_counter是否不支持DocumentSummaryIndex,是否需要使用原生token计数器?

代码示例

llm = MockLLM(max_tokens=256)
embed_model = MockEmbedding(embed_dim=1536)

token_counter = TokenCountingHandler(
    tokenizer=tiktoken.encoding_for_model("gpt-3.5-turbo").encode
)

callback_manager = CallbackManager([token_counter])

service_context = ServiceContext.from_defaults(
    llm=llm, 
    embed_model=embed_model, 
    callback_manager=callback_manager
)

response_synthesizer = get_response_synthesizer(
    summary_template = _chat_summarize_prompt(),  # TreeSummarize prompt
    text_qa_template = _chat_qa_prompt(),         # qa prompt
    response_mode    = "tree_summarize",
    use_async        = True,
)

doc_summary_index = DocumentSummaryIndex.from_documents(
    documents, # You need documents before execute this function
    response_synthesizer = response_synthesizer,
    service_context      = service_context,
    summary_query        = _summary_prompt(), # summary prompt
)

print(
"Embedding Tokens: ",
token_counter.total_embedding_token_count,
"\n",
"LLM Prompt Tokens: ",
token_counter.prompt_llm_token_count,
"\n",
"LLM Completion Tokens: ",
token_counter.completion_llm_token_count,
"\n",
"Total LLM Token Count: ",
token_counter.total_llm_token_count,
"\n",
)

# reset counts
token_counter.reset_counts()

问题原因及解决方案

  • 核心原因:Mock模型不触发回调
    TokenCountingHandler依赖LLM和Embedding模型的回调事件统计token,但MockLLM和MockEmbedding作为模拟实现,默认不会触发任何回调,导致计数器数值始终为0,和DocumentSummaryIndex本身无关。

  • 验证适配性:换用真实模型
    TokenCountingHandler完全支持DocumentSummaryIndex。如果替换为真实的LLM(如OpenAI的gpt-3.5-turbo)和Embedding模型(如text-embedding-ada-002),在创建索引生成摘要的过程中,回调会自动触发,token计数会正常更新。

  • Mock模型下的调试方案
    若要在Mock环境下验证计数逻辑,需手动给Mock模型添加回调触发代码:

    • 对MockLLM,在其complete或chat方法内,调用self.callback_manager.on_llm_prompt和self.callback_manager.on_llm_completion,传入计算好的prompt和completion token数。
    • 对MockEmbedding,在其get_text_embedding方法内,调用self.callback_manager.on_embed,传入文本的token数。
  • 原生计数器的必要性
    无需切换到原生计数器。只有当你需要脱离Llama Index的回调体系,做更定制化的token统计时,才需要直接使用tiktoken手动计算,但这会增加代码复杂度。

内容的提问来源于stack exchange,提问作者msxplus

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.09 15:46:21