为何Langchain的with_structured_output忽略Pydantic模型的可选参数?
Langchain结构化输出:Pydantic模型与JSON Schema的行为差异
问题背景
我在测试Langchain通过with_structured_output从LLM获取类JSON结构化输出的功能时,发现直接传入Pydantic模型和传入其JSON Schema的处理逻辑不一致,尤其是对Optional字段和默认值的处理,不符合预期。
测试代码
from typing import Optional from pydantic import BaseModel, Field from langchain_openai import ChatOpenAI class Joke(BaseModel): """Joke to tell user.""" setup: str = Field(description="The setup of the joke") punchline: str = Field(description="The punchline to the joke") rating: Optional[int] = Field(None, description="How funny the joke is, from 1 to 10") nonsense: Optional[str] = Field(None, description="Placeholder, always return None") class Joke2(BaseModel): """Joke to tell user.""" setup: str = Field(description="The setup of the joke") punchline: str = Field(description="The punchline to the joke") rating: int = Field(description="How funny the joke is, from 1 to 10") nonsense: str = Field(description="Placeholder, always return None") llm = ChatOpenAI(temperature=0, model_name="gpt-3.5-turbo")
测试结果与预期差异
我原本预期:
- 测试1和测试2的输出一致(返回
rating且nonsense为null) - 测试3和测试4的输出一致(返回
rating且nonsense为字符串'None')
但实际结果只有测试2和4符合预期,测试1和3出现偏差:
测试1:直接传入Joke模型
# 预期返回`rating`且`nonsense`为null llm.with_structured_output(Joke).invoke("Tell me a joke about cats") # 实际输出 {'setup': 'Why was the cat sitting on the computer?', 'punchline': 'To keep an eye on the mouse!'}
问题:缺失rating和nonsense字段
测试2:传入Joke的JSON Schema
# 预期返回`rating`且`nonsense`为null llm.with_structured_output(Joke.model_json_schema()).invoke("Tell me a joke about cats") # 实际输出 {'setup': 'Why was the cat sitting on the computer?', 'punchline': 'To keep an eye on the mouse!', 'rating': 8, 'nonsense': None}
符合预期
测试3:直接传入Joke2模型
# 预期返回`rating`且`nonsense`为字符串'None' llm.with_structured_output(Joke2).invoke("Tell me a joke about cats") # 实际输出 {'setup': 'Why was the cat sitting on the computer?', 'punchline': 'To keep an eye on the mouse!', 'rating': 5, 'nonsense': 'This joke is purr-fectly hilarious!'}
问题:nonsense未返回字符串'None',而是生成了无关内容
测试4:传入Joke2的JSON Schema
# 预期返回`rating`且`nonsense`为字符串'None' llm.with_structured_output(Joke2.model_json_schema()).invoke("Tell me a joke about cats") # 实际输出 {'setup': 'Why was the cat sitting on the computer?', 'punchline': 'To keep an eye on the mouse!', 'rating': 8, 'nonsense': 'None'}
符合预期
疑问
为什么直接传入Pydantic模型的测试1和3不符合预期,而传入JSON Schema的测试2和4却能达到预期效果?
内容的提问来源于stack exchange,提问作者Peter
相关产品推荐
相关产品推荐

