You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何排查两个Pandera Schema看似相同却不等的差异?

如何对比两个Pandera Schema的差异

当两个Pandera Schema的__repr__输出一致但实例判定不等时,可以通过以下方法定位差异:

1. 逐个字段对比列属性

Pandera的DataFrameSchema由多个Column对象组成,每个列包含dtype、nullable、checks、unique等属性,直接遍历对比这些属性可以精准定位差异:

import pandera as pa

schema1 = pa.infer_schema(df1)
schema2 = pa.infer_schema(df2)

# 检查字段数量与字段名差异
if len(schema1.columns) != len(schema2.columns):
    print(f"字段数量不一致:schema1({len(schema1.columns)}) vs schema2({len(schema2.columns)})")
    print(f"schema1独有字段:{set(schema1.columns) - set(schema2.columns)}")
    print(f"schema2独有字段:{set(schema2.columns) - set(schema1.columns)}")

# 对比共有字段的属性
for col_name in set(schema1.columns) & set(schema2.columns):
    col1 = schema1.columns[col_name]
    col2 = schema2.columns[col_name]
    
    if col1.dtype != col2.dtype:
        print(f"[{col_name}] dtype差异:{col1.dtype} vs {col2.dtype}")
    if col1.nullable != col2.nullable:
        print(f"[{col_name}] nullable设置差异:{col1.nullable} vs {col2.nullable}")
    if col1.unique != col2.unique:
        print(f"[{col_name}] unique设置差异:{col1.unique} vs {col2.unique}")
    if col1.coerce != col2.coerce:
        print(f"[{col_name}] coerce设置差异:{col1.coerce} vs {col2.coerce}")
    
    # 对比校验规则列表
    if len(col1.checks) != len(col2.checks):
        print(f"[{col_name}] 校验规则数量差异:{len(col1.checks)} vs {len(col2.checks)}")
    else:
        for idx, (check1, check2) in enumerate(zip(col1.checks, col2.checks)):
            if check1 != check2:
                print(f"[{col_name}] 第{idx+1}个校验规则差异:{check1} vs {check2}")

2. 转换为字典后递归对比

利用Schema的to_dict()方法将其转为字典结构,再通过递归对比字典的键值对,能快速找出深层差异:

# 转换为字典
schema_dict1 = schema1.to_dict()
schema_dict2 = schema2.to_dict()

# 递归对比字典的函数
def compare_schema_dicts(d1, d2, current_path=""):
    # 对比d1中的键
    for key in d1:
        if key not in d2:
            print(f"路径 {current_path}.{key} 仅存在于schema1")
        else:
            val1, val2 = d1[key], d2[key]
            if isinstance(val1, dict) and isinstance(val2, dict):
                compare_schema_dicts(val1, val2, f"{current_path}.{key}")
            elif isinstance(val1, list) and isinstance(val2, list):
                if len(val1) != len(val2):
                    print(f"路径 {current_path}.{key} 列表长度差异:{len(val1)} vs {len(val2)}")
                else:
                    for idx, (item1, item2) in enumerate(zip(val1, val2)):
                        if isinstance(item1, dict) and isinstance(item2, dict):
                            compare_schema_dicts(item1, item2, f"{current_path}.{key}[{idx}]")
                        elif item1 != item2:
                            print(f"路径 {current_path}.{key}[{idx}] 值差异:{item1} vs {item2}")
            elif val1 != val2:
                print(f"路径 {current_path}.{key} 值差异:{val1} vs {val2}")
    # 检查d2中独有的键
    for key in d2:
        if key not in d1:
            print(f"路径 {current_path}.{key} 仅存在于schema2")

compare_schema_dicts(schema_dict1, schema_dict2)

3. 检查Schema的全局属性

除了列属性,Schema本身的全局属性也可能导致不等,比如strict、metadata、name等:

if schema1.strict != schema2.strict:
    print(f"strict模式差异:schema1({schema1.strict}) vs schema2({schema2.strict})")
if schema1.metadata != schema2.metadata:
    print(f"元数据差异:schema1({schema1.metadata}) vs schema2({schema2.metadata})")
if schema1.name != schema2.name:
    print(f"Schema名称差异:schema1({schema1.name}) vs schema2({schema2.name})")

内容的提问来源于stack exchange,提问作者Galen

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.28 20:33:31