You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何预估算GPT-4 ChatCompletion的prompt_tokens数量

预估GPT-4 ChatCompletion接口的prompt_tokens数量

我用Node.js调用GPT-4的ChatCompletion接口时,返回的prompt_tokens为24。想在发起请求前用Python的tiktoken工具预估这个数值,但不清楚该如何将请求中的messages转换为对应文本传入脚本,才能让输出的token数和实际返回的24一致。

问题中的请求消息

Node.js代码里的消息内容如下:

  • System角色:You are a useful assistant.
  • User角色:What is the meaning of life?

正确的文本转换规则

OpenAI对ChatCompletion的消息采用特定格式编码token,需要将每个消息按照以下结构拼接:
<|im_start|>{role}\n{content}<|im_end|>
最后还要追加<|im_start|>assistant(模型需要为assistant的回复预留起始标记)。

对应到本次请求,转换后的完整文本为:

<|im_start|>system
You are a useful assistant.<|im_end|><|im_start|>user
What is the meaning of life?<|im_end|><|im_start|>assistant

验证脚本修改与测试

将上述文本传入tiktoken脚本,就能得到24的token计数。可以修改Python脚本直接使用该文本进行测试:

import tiktoken

# 转换后的完整文本
text = "<|im_start|>system\nYou are a useful assistant.<|im_end|><|im_start|>user\nWhat is the meaning of life?<|im_end|><|im_start|>assistant"
enc = tiktoken.encoding_for_model("gpt-4")
tokens = enc.encode(text)
token_count = len(tokens)
print(token_count)  # 输出:24

内容的提问来源于stack exchange,提问作者Shivam Sinha

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.25 12:17:19