如何让AutoGen多智能体输出符合预期的统一JSON格式?
AutoGen评审智能体JSON输出格式一致性解决方案
问题背景
使用AutoGen多智能体系统,配置了多个评审智能体对写作智能体生成的报告进行评审,要求评审结果以指定JSON格式输出,但当前输出格式混乱不一致,不符合预期。尝试在LLM配置中设置response_format = {"type": response_format}参数无效。
当前实际输出示例:
case1: {"heading":<heading>,"content":content} case2: [{"reviewer"::<reviewer>,review:{heading...
期望输出格式:
[ { "heading": "<评论1的标题>", "content": "<评论1的Markdown格式内容>" }, { "heading": "<评论2的标题>", "content": "<评论2的Markdown格式内容>" }, ... ]
可行解决方法
1. 强化系统提示词的格式约束
在评审智能体的系统提示词中,明确且强制要求输出格式,避免模糊表述,同时强调不得添加任何额外内容。示例提示词:
你是专业的报告评审员,需对给定报告进行评审并输出结构化结果。 **必须严格输出以下JSON格式的内容,不得包含任何JSON以外的文字、注释或说明**: [ { "heading": "<评论标题>", "content": "<Markdown格式的评论内容>" } ]
可在提示词中重复强调格式要求,或加入“若输出不符合格式,将视为无效评审”这类约束性语句,强化LLM的格式输出意识。
2. 使用AutoGen函数调用强制输出结构
通过定义接收评审结果的函数,让评审智能体通过调用该函数返回结果,利用LLM的工具调用特性强制输出符合参数结构的内容。
示例代码:
from autogen import AssistantAgent, UserProxyAgent # 定义接收评审结果的函数 def submit_review(reviews: list[dict]): """提交评审结果的函数,参数为包含heading和content的评论列表""" return reviews # 配置评审智能体 critic_agent = AssistantAgent( name="Critic", system_message="你是报告评审员,调用submit_review函数提交评审结果,参数必须符合要求的格式", llm_config={ "model": "gpt-4", "tools": [ { "type": "function", "function": { "name": "submit_review", "description": "提交评审结果", "parameters": { "type": "object", "properties": { "reviews": { "type": "array", "items": { "type": "object", "properties": { "heading": {"type": "string"}, "content": {"type": "string"} }, "required": ["heading", "content"] } } }, "required": ["reviews"] } } } ], "tool_choice": {"type": "function", "function": {"name": "submit_review"}} # 强制调用该函数 } ) # 用户代理用于触发评审 user_proxy = UserProxyAgent( name="UserProxy", code_execution_config={"work_dir": "coding"}, human_input_mode="NEVER" ) # 触发评审任务 user_proxy.initiate_chat(critic_agent, message="请评审这份报告:[报告内容]")
这种方式下,LLM会严格按照函数参数的JSON结构返回结果,避免格式混乱。
3. 后处理验证与自动修复
在评审智能体输出后,添加JSON格式验证步骤,若解析失败则触发智能体重新生成,或提取有效JSON内容进行修复。
示例代码:
import json import re def validate_and_fix_json(output): # 提取输出中的JSON部分(匹配最外层的[]或{}) json_match = re.search(r'\[.*?\]|\{.*?\}', output, re.DOTALL) if not json_match: return None json_str = json_match.group() try: # 尝试解析JSON return json.loads(json_str) except json.JSONDecodeError: # 若解析失败,返回None,触发重新生成 return None # 在智能体对话回调中处理输出 def check_review_output(sender, receiver, message): output = message["content"] valid_json = validate_and_fix_json(output) if not valid_json: # 通知评审智能体重新输出 receiver.send("你的输出格式不符合要求,请严格按照指定JSON格式重新输出", sender) else: # 处理有效结果 print("有效评审结果:", valid_json) # 给评审智能体添加回调 critic_agent.register_reply([UserProxyAgent, AssistantAgent], check_review_output, position=0)
通过这种方式,可以过滤无效输出,强制智能体修正格式。
4. 修正LLM的response_format配置
部分LLM(如OpenAI的GPT-4o、GPT-3.5-turbo-1106)支持response_format参数强制输出JSON,但需注意参数的正确写法:
llm_config = { "model": "gpt-4o", "response_format": {"type": "json_object"}, # 指定输出JSON对象 # 结合提示词明确要求输出数组结构 "system_message": "你是评审员,必须输出包含评论对象的JSON数组,格式如下:[{'heading': '', 'content': ''}]" }
注意:部分模型的response_format仅支持json_object,无法直接指定数组,此时需要在提示词中明确要求输出数组结构,结合参数约束实现预期格式。
内容的提问来源于stack exchange,提问作者Nayana Madhu
相关产品推荐
相关产品推荐

