如何用Pydantic根据字段条件校验JSON字段是否存在?
使用Pydantic实现带条件的JSON字段校验
需求说明
需要校验的JSON结构如下,字段的必填/禁止要求由type字段的值决定:
{ "id": "string", "type": "枚举值:initial/calc", "source": "仅当type为initial时必填", "source_path": "仅当type为initial时必填", "formula": "仅当type为calc时必填" }
现有代码问题
你尝试的字段级校验器(@validator)无法实现需求,因为它只能处理单个字段,无法关联其他字段的逻辑,也不能校验字段的存在性/冗余性。比如当type为calc时,传入source和source_path不会触发错误。
解决方案
方法1:根校验器(Root Validator)
用@root_validator可以拿到整个模型的所有字段值,从而实现跨字段的条件校验:
from enum import Enum from pydantic import BaseModel, root_validator, ValidationError class Type(Enum): initial = "initial" calc = "calc" class Tag(BaseModel): id: str type: Type # 先声明所有可能的字段,默认设为None source: str = None source_path: str = None formula: str = None @root_validator(pre=True) def check_fields_by_type(cls, values): tag_type = values.get("type") if not tag_type: return values if tag_type == Type.initial: # 校验必填字段 if "source" not in values or "source_path" not in values: raise ValueError("type为initial时,必须提供source和source_path") # 禁止冗余字段 if "formula" in values: raise ValueError("type为initial时,不能提供formula") elif tag_type == Type.calc: if "formula" not in values: raise ValueError("type为calc时,必须提供formula") if "source" in values or "source_path" in values: raise ValueError("type为calc时,不能提供source和source_path") return values
测试验证:
# 合法的calc类型数据 valid_calc = {"id": "tag_001", "type": "calc", "formula": "x*y+z"} Tag.parse_obj(valid_calc) # 正常通过 # 非法的calc类型(带source) invalid_calc = {"id": "tag_001", "type": "calc", "source": "local", "source_path": "/data"} try: Tag.parse_obj(invalid_calc) except ValidationError as e: print(e) # 会抛出错误提示 # 非法的initial类型(缺source_path) invalid_initial = {"id": "tag_002", "type": "initial", "source": "s3"} try: Tag.parse_obj(invalid_initial) except ValidationError as e: print(e) # 抛出必填字段缺失的错误
方法2:模型拆分继承(更优雅)
把不同类型的标签拆分成独立模型,利用Pydantic的类型判断自动校验,这种方式代码更清晰,扩展性更强。
Pydantic v2版本(推荐)
v2原生支持discriminator,可以自动根据type字段匹配对应的模型:
from enum import Enum from pydantic import BaseModel, Field, ValidationError class Type(Enum): initial = "initial" calc = "calc" class InitialTag(BaseModel): type: Type = Field(Type.initial, const=True) id: str source: str source_path: str class CalcTag(BaseModel): type: Type = Field(Type.calc, const=True) id: str formula: str class Tag(BaseModel): __root__: InitialTag | CalcTag = Field(..., discriminator="type")
使用示例:
# 合法数据 valid_data = {"id": "tag_003", "type": "initial", "source": "oss", "source_path": "/logs"} tag = Tag.model_validate(valid_data) print(tag.__root__) # 输出InitialTag实例 # 非法数据(calc类型带source) invalid_data = {"id": "tag_004", "type": "calc", "source": "local"} try: Tag.model_validate(invalid_data) except ValidationError as e: print(e) # 提示存在额外字段source
Pydantic v1版本
v1可以用Union结合手动解析:
from enum import Enum from pydantic import BaseModel, Field, ValidationError from typing import Union class Type(Enum): initial = "initial" calc = "calc" class InitialTag(BaseModel): type: Type = Field(Type.initial, const=True) id: str source: str source_path: str class CalcTag(BaseModel): type: Type = Field(Type.calc, const=True) id: str formula: str Tag = Union[InitialTag, CalcTag] def parse_tag(data): try: return InitialTag.parse_obj(data) except ValidationError: return CalcTag.parse_obj(data)
是否需要自定义解析函数?
完全不需要,Pydantic完全能覆盖这类条件校验场景。上面的两种方法都能稳定实现需求,其中模型拆分的方式更符合Pydantic的设计思路,代码可读性和可维护性更高。
内容的提问来源于stack exchange,提问作者Motixa
相关产品推荐
相关产品推荐

