You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Databricks自定义Logger无法读取extra参数问题排查求助

问题排查:Python Logging中extra参数无法在自定义Handler中获取的原因与解决方法

问题核心原因

Python标准库logging模块对extra参数的处理逻辑是:将传入的字典中的每个键值对直接添加到LogRecord对象的属性中,而非把整个字典存储在名为extra的属性里。你代码里的log_entry.__dict__.get("extra", None)永远返回None,因为根本不存在这个键。

举个例子:你调用logger.debug("msg", extra={"Test":"testvalue"}),最终record.Test会直接等于"testvalue",而不是record.extra["Test"]。

修复方案

方案1:直接提取自定义extra字段

可以通过过滤logging默认字段的方式,从record中提取所有额外传入的字段:

修改emit方法的相关代码:

def emit(self, record):
    # 转换时间戳为ISO格式
    iso_timestamp = datetime.utcfromtimestamp(record.created).isoformat()
    
    # 定义logging默认字段集合,用于过滤
    default_log_fields = {
        'args', 'asctime', 'created', 'exc_info', 'exc_text', 'filename',
        'funcName', 'levelname', 'levelno', 'lineno', 'module', 'msecs',
        'message', 'msg', 'name', 'pathname', 'process', 'processName',
        'relativeCreated', 'stack_info', 'thread', 'threadName'
    }

    log_data = {
        "timestamp": iso_timestamp,
        "message": record.msg,
    }

    # 遍历record属性,提取非默认字段(即extra传入的内容)
    for key, value in record.__dict__.items():
        if key not in default_log_fields:
            if isinstance(value, str):
                log_data[key] = value
            else:
                log_data[f'{key}_json'] = json.dumps(value)

    log_json = json.dumps(log_data, default=str)
    self.upload_to_azure_datalake(log_json)

方案2:约定extra字段的统一存储方式(需修改日志调用)

如果希望把整个extra字典作为整体处理,可以在调用日志时,将extra内容放在一个特定键下:

# 日志调用时调整写法
logger.debug("This is a debug message.", extra={"extra_fields": {"Test":"testvalue"}})

然后在emit方法中直接获取该键:

def emit(self, record):
    iso_timestamp = datetime.utcfromtimestamp(record.created).isoformat()
    log_data = {
        "timestamp": iso_timestamp,
        "message": record.msg,
    }

    extra_fields = record.__dict__.get("extra_fields", None)
    if extra_fields is not None:
        for key, value in extra_fields.items():
            if isinstance(value, str):
                log_data[key] = value
            else:
                log_data[f'{key}_json'] = json.dumps(value)

    log_json = json.dumps(log_data, default=str)
    self.upload_to_azure_datalake(log_json)

额外注意点

你的upload_to_azure_datalake方法使用了overwrite=True,这会导致每次写入日志都覆盖之前的文件内容,最终只会保留最后一条日志。建议改成追加模式:

def upload_to_azure_datalake(self, log_json):
    file_path = f"{self.directory_name}/{self.file_name}"
    file_client = self.file_system_client.get_file_client(file_path)

    if not file_client.exists():
        file_client.create_file()
        file_client.upload_data(log_json + "\n", overwrite=True)
    else:
        # 追加日志到文件末尾
        current_length = file_client.get_file_properties().size
        file_client.append_data(log_json + "\n", offset=current_length, length=len(log_json + "\n"))
        file_client.flush_data(current_length + len(log_json + "\n"))

内容的提问来源于stack exchange,提问作者Bee_Riii

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.04 02:03:14