You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何分别跟踪调用同一Azure部署OpenAI GPT模型的程序开销?

如何分别统计两个程序调用同一Azure OpenAI部署的开销

方法一:利用Azure自带的监控与成本分析功能

  • 添加自定义请求标签:在调用openai.ChatCompletion.create()时,通过extra_body参数添加程序标识标签,Azure会将标签关联到监控数据中,实现后续按程序拆分统计。
    示例代码修改(程序A):
    response = openai.ChatCompletion.create(
        engine="gpt-4-32k-viet",
        messages=conversation,
        temperature=.7,
        max_tokens=max_response_tokens,
        extra_body={"tags": {"program": "A"}}  # 程序B改为"program: B"
    )
    
  • Azure Monitor筛选统计:进入Azure OpenAI资源的「监控」面板,选择「Token Count」指标,添加筛选条件为自定义标签program,即可分别查看两个程序的输入/输出token消耗,结合模型单价计算开销。
  • 成本管理细化分析:在Azure「成本管理 + 计费」中创建报表,将自定义标签设为分组依据或筛选条件,直接拆分两个程序的费用。

方法二:代码层面手动统计token并核算费用

基于你提供的代码,扩展token统计逻辑,记录每个程序的token消耗并计算费用:

  1. 统计单次请求token:输入token通过num_tokens_from_messages()获取,输出token从API响应的usage字段提取。
  2. 记录与汇总:将每次请求的token数和计算出的费用写入日志或存储,定期汇总得到总开销。

修改后的代码示例(程序A):

import tiktoken
import openai
import os
from datetime import datetime

openai.api_type = "azure"
openai.api_version = "2023-03-15-preview"
openai.api_base = "https://[resourcename].openai.azure.com/"
openai.api_key = "[my instance key]"

# 模型单价(需按Azure实际定价调整)
INPUT_TOKEN_PRICE = 0.06 / 1000000  # GPT-4 32k输入token:$0.06/百万
OUTPUT_TOKEN_PRICE = 0.12 / 1000000  # GPT-4 32k输出token:$0.12/百万

system_message = {"role": "system", "content": "You are a helpful assistant."}
max_response_tokens = 250
token_limit= 4096
conversation=[]
conversation.append(system_message)

def num_tokens_from_messages(messages, model="gpt-4-32k"):
    encoding = tiktoken.encoding_for_model(model)
    num_tokens = 0
    for message in messages:
        num_tokens += 4  # 每个消息的固定token开销
        for key, value in message.items():
            num_tokens += len(encoding.encode(value))
            if key == "name":
                num_tokens += -1  # 有name时role的token会被省略
    num_tokens += 2  # 回复的固定前缀token
    return num_tokens

def log_cost(input_tokens, output_tokens, program_name):
    # 计算单次请求费用
    cost = input_tokens * INPUT_TOKEN_PRICE + output_tokens * OUTPUT_TOKEN_PRICE
    # 写入日志(可替换为写入数据库)
    with open(f"{program_name}_cost_log.txt", "a", encoding="utf-8") as f:
        log_line = f"{datetime.now()},输入token:{input_tokens},输出token:{output_tokens},费用:{cost:.6f}$\n"
        f.write(log_line)

user_input = 'Hi there. What is the difference between Facebook and TikTok?'
conversation.append({"role": "user", "content": user_input})
conv_history_tokens = num_tokens_from_messages(conversation)

while (conv_history_tokens + max_response_tokens >= token_limit):
    del conversation[1]
    conv_history_tokens = num_tokens_from_messages(conversation)

response = openai.ChatCompletion.create(
    engine="gpt-4-32k-viet",
    messages=conversation,
    temperature=.7,
    max_tokens=max_response_tokens,
)

# 获取输出token数并记录费用
output_tokens = response['usage']['completion_tokens']
log_cost(conv_history_tokens, output_tokens, "program_A")

conversation.append({"role": "assistant", "content": response['choices'][0]['message']['content']})
print("\n" + response['choices'][0]['message']['content'] + "\n")

程序B只需将log_cost中的program_A改为program_B,后续通过各自日志文件汇总总开销。

方法三:使用不同API密钥区分

在Azure OpenAI资源的「密钥和终结点」页面创建两个独立API密钥,分别分配给程序A和程序B。之后在Azure成本管理中,通过「API密钥」维度筛选,即可拆分两个程序的费用。


内容的提问来源于stack exchange,提问作者Franck Dernoncourt

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.22 22:57:06