如何利用LLM或NLP模型识别用户指令中的单/多任务?
解决单/多任务指令区分问题的方案
一、修复当前GPT-3.5提示词的问题
你的核心问题是提示词对「多任务」的定义覆盖不全,且示例场景单一,导致模型误判带触发条件的多任务指令。可以从以下两点优化提示词:
明确多任务判定规则
在提示词开头补充清晰的判定标准:- 单任务:仅包含一个需要执行的动作(无论是否带前置触发条件)
- 多任务:包含两个及以上并列的执行动作,共享同一个触发条件或存在多个独立动作指令
补充对应场景的示例
加入你遇到的触发式多任务指令作为正确示例,让模型学习这类场景的判定逻辑。
优化后的提示词示例:
Carefully analyze the request to determine if it contains a single task or multiple tasks, then generate JSON output following these rules: - **Single Task**: The request contains only one action to execute (regardless of whether it has a trigger condition). - **Multiple Task**: The request contains two or more parallel actions to execute, either sharing the same trigger condition or being independent tasks. ### Example 1 (Query-type Multiple Task): User Request: "how many executions happen with success and fail so far" Analysis: This is a multiple task because it asks for two distinct counts: successful executions and failed executions. JSON Output: { "requests": [ { "task": "how many executions happen with success and fail so far", "type": "multiple_task" } ] } ### Example 2 (Trigger-based Multiple Task): User Request: "When a New google calendar event is created, post a message to general channel in slack plus sync it to Salesforce leads" Analysis: This is a multiple task because it requires executing two parallel actions (post to Slack, sync to Salesforce) under the same trigger condition (new Google Calendar event created). JSON Output: { "requests": [ { "task": "When a New google calendar event is created, post a message to general channel in slack plus sync it to Salesforce leads", "type": "multiple_task" } ] } ### Example 3 (Single Task): User Request: "Post a message to the general Slack channel" Analysis: This is a single task because it only contains one action. JSON Output: { "requests": [ { "task": "Post a message to the general Slack channel", "type": "single_task" } ] } Now process the following user request: User request:
二、提示词工程vs微调的抉择建议
优先选择提示词工程
针对指令表述多样的问题,可通过场景化示例覆盖解决:收集业务中常见的指令类型(触发式、查询式、命令式等),为每种类型补充单/多任务的标注示例,让模型学习你的业务判定逻辑。这种方式成本低、迭代快,无需大规模数据集,适合初期快速验证。提示词达瓶颈时再考虑微调
如果提示词优化后仍有大量边缘场景误判,再启动微调。此时无需构建大规模数据集,只需:- 收集业务中高频误判、边缘场景的指令
- 为每条指令标注正确的「task」和「type」
- 使用LoRA等轻量微调方法,基于GPT-3.5基础模型做针对性微调,既能降低计算成本,又能精准适配业务场景的判定需求
内容的提问来源于stack exchange,提问作者Rishabh Tripathi
相关产品推荐
相关产品推荐

