You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何自定义pandas_profiling/y_data_profiling的告警及其他指标

自定义y_data_profiling(原pandas_profiling)的告警与指标

一、自定义告警规则

y_data_profiling支持通过CustomWarning类添加完全自定义的告警逻辑,你可以针对单列、多列甚至整个数据集设置告警条件。

示例:针对某列的自定义阈值告警

比如要给age列设置告警:当列中超过90的值占比超过5%时触发告警。

import pandas as pd
from ydata_profiling import ProfileReport
from ydata_profiling.config import Settings
from ydata_profiling.model.warning import CustomWarning

# 构造测试数据
df = pd.DataFrame({
    "age": [18, 25, 95, 92, 30]*20 + [100]*10,  # 10个100,占比10%
    "income": [5000, 6000, 8000]*36
})

# 定义告警逻辑函数
def age_over_90_warning(series):
    over_90_ratio = (series > 90).mean()
    if over_90_ratio > 0.05:
        return f"超过90的数值占比{over_90_ratio:.2%},超过5%阈值"
    return None

# 创建自定义告警实例
custom_warnings = [
    CustomWarning(
        column="age",
        name="age_over_90",
        message_func=age_over_90_warning,
        category="数值异常"
    )
]

# 生成报告时传入自定义告警
profile = ProfileReport(
    df,
    settings=Settings(
        warnings={"custom_warnings": custom_warnings}
    )
)
profile.to_file("custom_warning_report.html")

示例:数据集级别的自定义告警

比如检测数据集的重复行占比超过1%时触发告警:

def duplicate_rows_warning(df):
    duplicate_ratio = df.duplicated().mean()
    if duplicate_ratio > 0.01:
        return f"重复行占比{duplicate_ratio:.2%},超过1%阈值"
    return None

custom_warnings = [
    CustomWarning(
        column=None,  # None表示数据集级别
        name="duplicate_rows",
        message_func=duplicate_rows_warning,
        category="数据重复"
    )
]

二、添加自定义指标

你可以通过继承CustomUnivariateMetric(单变量指标)或CustomMultivariateMetric(多变量指标)类来添加新的统计指标,并将其整合到报告的对应模块中。

示例:添加单列的「中位数与均值差值」指标

from ydata_profiling.model.metrics import CustomUnivariateMetric
from ydata_profiling.report.presentation.core import VariableInfo
from ydata_profiling.config import Settings

# 定义自定义单变量指标类
class MeanMedianDiffMetric(CustomUnivariateMetric):
    def calculate(self, series, settings: Settings):
        # 只针对数值型列计算
        if pd.api.types.is_numeric_dtype(series):
            mean_val = series.mean()
            median_val = series.median()
            return abs(mean_val - median_val)
        return None

    def get_presentation(self, value, variable):
        # 定义指标在报告中的展示方式
        return VariableInfo(
            name="均值中位数差值",
            value=f"{value:.2f}" if value is not None else "N/A",
            anchor_id=f"{variable.name}_mean_median_diff"
        )

# 注册自定义指标
custom_metrics = {
    "mean_median_diff": MeanMedianDiffMetric()
}

# 生成报告时传入自定义指标
profile = ProfileReport(
    df,
    settings=Settings(
        variables={"num": {"custom_metrics": custom_metrics}}
    )
)
profile.to_file("custom_metric_report.html")

这个指标会出现在数值型列的「统计指标」模块中,展示该列均值与中位数的绝对差值。

三、修改默认告警阈值

如果只是想调整现有告警的触发条件,不需要新增逻辑,可以直接修改配置中的阈值参数:

profile = ProfileReport(
    df,
    settings=Settings(
        warnings={
            "thresholds": {
                "missing": 0.1,  # 缺失值占比超过10%触发告警(默认是0.5)
                "correlation": 0.9,  # 相关系数超过0.9触发告警(默认是0.9)
                "duplicates": 0.05  # 重复行占比超过5%触发告警(默认是0.1)
            }
        }
    )
)

内容的提问来源于stack exchange,提问作者wantering_otter

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.22 09:55:07