You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

为何Langchain的with_structured_output忽略Pydantic模型的可选参数?

Langchain结构化输出:Pydantic模型与JSON Schema的行为差异

问题背景

我在测试Langchain通过with_structured_output从LLM获取类JSON结构化输出的功能时,发现直接传入Pydantic模型和传入其JSON Schema的处理逻辑不一致,尤其是对Optional字段和默认值的处理,不符合预期。

测试代码

from typing import Optional
from pydantic import BaseModel, Field
from langchain_openai import ChatOpenAI

class Joke(BaseModel):
    """Joke to tell user."""

    setup: str = Field(description="The setup of the joke")
    punchline: str = Field(description="The punchline to the joke")
    rating: Optional[int] = Field(None, description="How funny the joke is, from 1 to 10")
    nonsense: Optional[str] = Field(None, description="Placeholder, always return None")

class Joke2(BaseModel):
    """Joke to tell user."""

    setup: str = Field(description="The setup of the joke")
    punchline: str = Field(description="The punchline to the joke")
    rating: int = Field(description="How funny the joke is, from 1 to 10")
    nonsense: str = Field(description="Placeholder, always return None")

llm = ChatOpenAI(temperature=0, model_name="gpt-3.5-turbo")

测试结果与预期差异

我原本预期:

  • 测试1和测试2的输出一致(返回rating且nonsense为null)
  • 测试3和测试4的输出一致(返回rating且nonsense为字符串'None')

但实际结果只有测试2和4符合预期,测试1和3出现偏差:

测试1:直接传入Joke模型

# 预期返回`rating`且`nonsense`为null
llm.with_structured_output(Joke).invoke("Tell me a joke about cats")
# 实际输出
{'setup': 'Why was the cat sitting on the computer?',
 'punchline': 'To keep an eye on the mouse!'}

问题:缺失rating和nonsense字段

测试2:传入Joke的JSON Schema

# 预期返回`rating`且`nonsense`为null
llm.with_structured_output(Joke.model_json_schema()).invoke("Tell me a joke about cats")
# 实际输出
{'setup': 'Why was the cat sitting on the computer?',
 'punchline': 'To keep an eye on the mouse!',
 'rating': 8,
 'nonsense': None}

符合预期

测试3:直接传入Joke2模型

# 预期返回`rating`且`nonsense`为字符串'None'
llm.with_structured_output(Joke2).invoke("Tell me a joke about cats")
# 实际输出
{'setup': 'Why was the cat sitting on the computer?',
 'punchline': 'To keep an eye on the mouse!',
 'rating': 5,
 'nonsense': 'This joke is purr-fectly hilarious!'}

问题:nonsense未返回字符串'None',而是生成了无关内容

测试4:传入Joke2的JSON Schema

# 预期返回`rating`且`nonsense`为字符串'None'
llm.with_structured_output(Joke2.model_json_schema()).invoke("Tell me a joke about cats")
# 实际输出
{'setup': 'Why was the cat sitting on the computer?',
 'punchline': 'To keep an eye on the mouse!',
 'rating': 8,
 'nonsense': 'None'}

符合预期

疑问

为什么直接传入Pydantic模型的测试1和3不符合预期,而传入JSON Schema的测试2和4却能达到预期效果?


内容的提问来源于stack exchange,提问作者Peter

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.22 15:42:48