You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Pydantic根据字段条件校验JSON字段是否存在?

使用Pydantic实现带条件的JSON字段校验

需求说明

需要校验的JSON结构如下,字段的必填/禁止要求由type字段的值决定:

{
    "id": "string",
    "type": "枚举值:initial/calc",
    "source": "仅当type为initial时必填",
    "source_path": "仅当type为initial时必填",
    "formula": "仅当type为calc时必填"
}

现有代码问题

你尝试的字段级校验器(@validator)无法实现需求,因为它只能处理单个字段,无法关联其他字段的逻辑,也不能校验字段的存在性/冗余性。比如当type为calc时,传入source和source_path不会触发错误。

解决方案

方法1:根校验器(Root Validator)

用@root_validator可以拿到整个模型的所有字段值,从而实现跨字段的条件校验:

from enum import Enum
from pydantic import BaseModel, root_validator, ValidationError


class Type(Enum):
    initial = "initial"
    calc = "calc"


class Tag(BaseModel):
    id: str
    type: Type
    # 先声明所有可能的字段,默认设为None
    source: str = None
    source_path: str = None
    formula: str = None

    @root_validator(pre=True)
    def check_fields_by_type(cls, values):
        tag_type = values.get("type")
        if not tag_type:
            return values
        
        if tag_type == Type.initial:
            # 校验必填字段
            if "source" not in values or "source_path" not in values:
                raise ValueError("type为initial时,必须提供source和source_path")
            # 禁止冗余字段
            if "formula" in values:
                raise ValueError("type为initial时,不能提供formula")
        elif tag_type == Type.calc:
            if "formula" not in values:
                raise ValueError("type为calc时,必须提供formula")
            if "source" in values or "source_path" in values:
                raise ValueError("type为calc时,不能提供source和source_path")
        return values

测试验证:

# 合法的calc类型数据
valid_calc = {"id": "tag_001", "type": "calc", "formula": "x*y+z"}
Tag.parse_obj(valid_calc)  # 正常通过

# 非法的calc类型(带source)
invalid_calc = {"id": "tag_001", "type": "calc", "source": "local", "source_path": "/data"}
try:
    Tag.parse_obj(invalid_calc)
except ValidationError as e:
    print(e)  # 会抛出错误提示

# 非法的initial类型(缺source_path)
invalid_initial = {"id": "tag_002", "type": "initial", "source": "s3"}
try:
    Tag.parse_obj(invalid_initial)
except ValidationError as e:
    print(e)  # 抛出必填字段缺失的错误

方法2:模型拆分继承(更优雅)

把不同类型的标签拆分成独立模型,利用Pydantic的类型判断自动校验,这种方式代码更清晰,扩展性更强。

Pydantic v2版本(推荐)

v2原生支持discriminator,可以自动根据type字段匹配对应的模型:

from enum import Enum
from pydantic import BaseModel, Field, ValidationError


class Type(Enum):
    initial = "initial"
    calc = "calc"


class InitialTag(BaseModel):
    type: Type = Field(Type.initial, const=True)
    id: str
    source: str
    source_path: str


class CalcTag(BaseModel):
    type: Type = Field(Type.calc, const=True)
    id: str
    formula: str


class Tag(BaseModel):
    __root__: InitialTag | CalcTag = Field(..., discriminator="type")

使用示例:

# 合法数据
valid_data = {"id": "tag_003", "type": "initial", "source": "oss", "source_path": "/logs"}
tag = Tag.model_validate(valid_data)
print(tag.__root__)  # 输出InitialTag实例

# 非法数据(calc类型带source)
invalid_data = {"id": "tag_004", "type": "calc", "source": "local"}
try:
    Tag.model_validate(invalid_data)
except ValidationError as e:
    print(e)  # 提示存在额外字段source

Pydantic v1版本

v1可以用Union结合手动解析:

from enum import Enum
from pydantic import BaseModel, Field, ValidationError
from typing import Union


class Type(Enum):
    initial = "initial"
    calc = "calc"


class InitialTag(BaseModel):
    type: Type = Field(Type.initial, const=True)
    id: str
    source: str
    source_path: str


class CalcTag(BaseModel):
    type: Type = Field(Type.calc, const=True)
    id: str
    formula: str


Tag = Union[InitialTag, CalcTag]

def parse_tag(data):
    try:
        return InitialTag.parse_obj(data)
    except ValidationError:
        return CalcTag.parse_obj(data)

是否需要自定义解析函数?

完全不需要,Pydantic完全能覆盖这类条件校验场景。上面的两种方法都能稳定实现需求,其中模型拆分的方式更符合Pydantic的设计思路,代码可读性和可维护性更高。

内容的提问来源于stack exchange,提问作者Motixa

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.07 04:20:15