关于Mistral v3 Tekken分词器聊天补全请求的技术问询
Mistral v3 Tekken分词器聊天补全编码疑问解析
实验代码与输出
tokenizer = MistralTokenizer.v3(is_tekken=True) chat_complete_req = ChatCompletionRequest( messages=[ SystemMessage(content="You are an agent working for an airlines."), AssistantMessage(content="Hi! How can I help you?"), UserMessage(content="What's the weather like today in Paris"), ] ) tokenizer.encode_chat_completion(chat_complete_req)
分词输出
Tokenized(tokens=[1, 3, 4, 37133, 1033, 3075, 1710, 1362, 3508, 1636, 1063, 2, 3, 4568, 1584, 1420, 10496, 5564, 1394, 1420, 110293, 1338, 7493, 1681, 1278, 17253, 2479, 9406, 1294, 6993, 4], text="<s>[INST][/INST]Hi! How can I help you?<s>[INST]You are an agent working for an airlines.\n\nWhat's the weather like today in Paris[/INST]", prefix_ids=None, images=[])
疑问解析
1. 分词后消息顺序为何发生改变?
Mistral v3 Tekken分词器的encode_chat_completion方法遵循模型训练时的对话格式要求,将助手消息作为历史上下文放在前面,把需要模型响应的用户查询(附带系统提示)放在最后一个[INST]块内,以此确保模型能正确识别对话上下文并生成符合预期的回复。
2. 系统消息与用户消息为何被合并?官方文档提及建议合并两者,这是默认合并的原因吗?
是的,这就是默认合并的原因。Mistral的对话格式规范明确要求,系统提示需与用户当前查询合并到同一个[INST]块内,用换行分隔。这种设计是为了让模型将系统指令作为约束,应用到当前用户查询的处理中,完全匹配模型预训练时的输入格式预期,避免因格式不匹配导致输出异常。
3. 助手消息前为何出现空的[INST][/INST]块?
这是Mistral Tekken格式的特殊设计:助手消息必须放置在[INST][/INST]块之后。当前输入中的助手消息属于历史对话上下文,需要通过空的[INST][/INST]块占位,来匹配模型训练时的历史对话结构,确保模型能正确解析上下文逻辑。
内容的提问来源于stack exchange,提问作者Dhineshkumar
相关产品推荐
相关产品推荐

