You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandera仅校验首个不满足项,如何实现全条件校验?

解决Pandera仅校验首个不符合条件的问题

Pandera默认是快速失败模式,碰到第一个校验错误就停止,要让它一次性校验所有条件并收集所有错误,只需要启用**延迟校验(lazy validation)**模式。

修改方案

在调用schema.validate()时添加lazy=True参数,同时捕获pa.errors.SchemaErrors(注意是复数形式)异常,这个异常会包含所有校验失败的案例。

修改后的完整代码

import pandas as pd
import pandera as pa

# Sample DataFrame
data = {
    "column1": [None, 2, 3, 4],
    "column2": ["A", None, "C", "D"],
    "column3": [46.0, 50.0, 30.0, 70.0],
}
df = pd.DataFrame(data)

# Define the validation schema
schema = pa.DataFrameSchema({
    "column1": pa.Column(int),
    "column2": pa.Column(str),
    "column3": pa.Column(float, checks=[
        pa.Check(lambda x: x is not None, error="shouldn't be None"),
        pa.Check(lambda x: x > 45.0, error="has to be > 45.0"),
    ]),
})

try:
    schema.validate(df, lazy=True)  # 启用lazy模式
    print("Validation successful!")
except pa.errors.SchemaErrors as e:  # 捕获复数形式的SchemaErrors
    print("Validation failed, all errors:")
    # 打印所有错误详情
    for err in e.errors:
        print(f"- {err}")
    
    # 提取所有失败的索引(去重,避免同一行因多错误重复过滤)
    all_failure_indices = e.failure_cases['index'].unique()
    clean_df = df[~df.index.isin(all_failure_indices)]
    not_clean_df = df[df.index.isin(all_failure_indices)]

print('\nclean_df\n', clean_df)
print('\nnot_clean_df\n', not_clean_df)

运行结果

Validation failed, all errors:
- non-nullable series 'column1' contains null values:
0   NaN
Name: column1, dtype: float64
- non-nullable series 'column2' contains null values:
1    None
Name: column2, dtype: object
- <lambda> check raised: has to be > 45.0
failure cases:
   index  column3
2      2     30.0

clean_df
    column1 column2  column3
3      4.0       D     70.0

not_clean_df
    column1 column2  column3
0      NaN       A     46.0
1      2.0    None     50.0
2      3.0       C     30.0

关键说明

  • lazy=True:让Pandera不中断校验,直到所有规则检查完毕,收集全部失败案例。
  • SchemaErrors(复数):该异常对象包含errors列表(每个错误的详细描述)和failure_cases DataFrame(所有失败行的索引、对应列及错误原因)。
  • 索引去重:同一行可能违反多个校验规则,用unique()去重后再过滤,避免重复处理同一行。

内容的提问来源于stack exchange,提问作者Bartosz Lewiński

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.25 19:04:57