如何自定义pandas_profiling/y_data_profiling的告警及其他指标
自定义y_data_profiling(原pandas_profiling)的告警与指标
一、自定义告警规则
y_data_profiling支持通过CustomWarning类添加完全自定义的告警逻辑,你可以针对单列、多列甚至整个数据集设置告警条件。
示例:针对某列的自定义阈值告警
比如要给age列设置告警:当列中超过90的值占比超过5%时触发告警。
import pandas as pd from ydata_profiling import ProfileReport from ydata_profiling.config import Settings from ydata_profiling.model.warning import CustomWarning # 构造测试数据 df = pd.DataFrame({ "age": [18, 25, 95, 92, 30]*20 + [100]*10, # 10个100,占比10% "income": [5000, 6000, 8000]*36 }) # 定义告警逻辑函数 def age_over_90_warning(series): over_90_ratio = (series > 90).mean() if over_90_ratio > 0.05: return f"超过90的数值占比{over_90_ratio:.2%},超过5%阈值" return None # 创建自定义告警实例 custom_warnings = [ CustomWarning( column="age", name="age_over_90", message_func=age_over_90_warning, category="数值异常" ) ] # 生成报告时传入自定义告警 profile = ProfileReport( df, settings=Settings( warnings={"custom_warnings": custom_warnings} ) ) profile.to_file("custom_warning_report.html")
示例:数据集级别的自定义告警
比如检测数据集的重复行占比超过1%时触发告警:
def duplicate_rows_warning(df): duplicate_ratio = df.duplicated().mean() if duplicate_ratio > 0.01: return f"重复行占比{duplicate_ratio:.2%},超过1%阈值" return None custom_warnings = [ CustomWarning( column=None, # None表示数据集级别 name="duplicate_rows", message_func=duplicate_rows_warning, category="数据重复" ) ]
二、添加自定义指标
你可以通过继承CustomUnivariateMetric(单变量指标)或CustomMultivariateMetric(多变量指标)类来添加新的统计指标,并将其整合到报告的对应模块中。
示例:添加单列的「中位数与均值差值」指标
from ydata_profiling.model.metrics import CustomUnivariateMetric from ydata_profiling.report.presentation.core import VariableInfo from ydata_profiling.config import Settings # 定义自定义单变量指标类 class MeanMedianDiffMetric(CustomUnivariateMetric): def calculate(self, series, settings: Settings): # 只针对数值型列计算 if pd.api.types.is_numeric_dtype(series): mean_val = series.mean() median_val = series.median() return abs(mean_val - median_val) return None def get_presentation(self, value, variable): # 定义指标在报告中的展示方式 return VariableInfo( name="均值中位数差值", value=f"{value:.2f}" if value is not None else "N/A", anchor_id=f"{variable.name}_mean_median_diff" ) # 注册自定义指标 custom_metrics = { "mean_median_diff": MeanMedianDiffMetric() } # 生成报告时传入自定义指标 profile = ProfileReport( df, settings=Settings( variables={"num": {"custom_metrics": custom_metrics}} ) ) profile.to_file("custom_metric_report.html")
这个指标会出现在数值型列的「统计指标」模块中,展示该列均值与中位数的绝对差值。
三、修改默认告警阈值
如果只是想调整现有告警的触发条件,不需要新增逻辑,可以直接修改配置中的阈值参数:
profile = ProfileReport( df, settings=Settings( warnings={ "thresholds": { "missing": 0.1, # 缺失值占比超过10%触发告警(默认是0.5) "correlation": 0.9, # 相关系数超过0.9触发告警(默认是0.9) "duplicates": 0.05 # 重复行占比超过5%触发告警(默认是0.1) } } ) )
内容的提问来源于stack exchange,提问作者wantering_otter
相关产品推荐
相关产品推荐

