Databricks自定义Logger无法读取extra参数问题排查求助
问题排查:Python Logging中extra参数无法在自定义Handler中获取的原因与解决方法
问题核心原因
Python标准库logging模块对extra参数的处理逻辑是:将传入的字典中的每个键值对直接添加到LogRecord对象的属性中,而非把整个字典存储在名为extra的属性里。你代码里的log_entry.__dict__.get("extra", None)永远返回None,因为根本不存在这个键。
举个例子:你调用logger.debug("msg", extra={"Test":"testvalue"}),最终record.Test会直接等于"testvalue",而不是record.extra["Test"]。
修复方案
方案1:直接提取自定义extra字段
可以通过过滤logging默认字段的方式,从record中提取所有额外传入的字段:
修改emit方法的相关代码:
def emit(self, record): # 转换时间戳为ISO格式 iso_timestamp = datetime.utcfromtimestamp(record.created).isoformat() # 定义logging默认字段集合,用于过滤 default_log_fields = { 'args', 'asctime', 'created', 'exc_info', 'exc_text', 'filename', 'funcName', 'levelname', 'levelno', 'lineno', 'module', 'msecs', 'message', 'msg', 'name', 'pathname', 'process', 'processName', 'relativeCreated', 'stack_info', 'thread', 'threadName' } log_data = { "timestamp": iso_timestamp, "message": record.msg, } # 遍历record属性,提取非默认字段(即extra传入的内容) for key, value in record.__dict__.items(): if key not in default_log_fields: if isinstance(value, str): log_data[key] = value else: log_data[f'{key}_json'] = json.dumps(value) log_json = json.dumps(log_data, default=str) self.upload_to_azure_datalake(log_json)
方案2:约定extra字段的统一存储方式(需修改日志调用)
如果希望把整个extra字典作为整体处理,可以在调用日志时,将extra内容放在一个特定键下:
# 日志调用时调整写法 logger.debug("This is a debug message.", extra={"extra_fields": {"Test":"testvalue"}})
然后在emit方法中直接获取该键:
def emit(self, record): iso_timestamp = datetime.utcfromtimestamp(record.created).isoformat() log_data = { "timestamp": iso_timestamp, "message": record.msg, } extra_fields = record.__dict__.get("extra_fields", None) if extra_fields is not None: for key, value in extra_fields.items(): if isinstance(value, str): log_data[key] = value else: log_data[f'{key}_json'] = json.dumps(value) log_json = json.dumps(log_data, default=str) self.upload_to_azure_datalake(log_json)
额外注意点
你的upload_to_azure_datalake方法使用了overwrite=True,这会导致每次写入日志都覆盖之前的文件内容,最终只会保留最后一条日志。建议改成追加模式:
def upload_to_azure_datalake(self, log_json): file_path = f"{self.directory_name}/{self.file_name}" file_client = self.file_system_client.get_file_client(file_path) if not file_client.exists(): file_client.create_file() file_client.upload_data(log_json + "\n", overwrite=True) else: # 追加日志到文件末尾 current_length = file_client.get_file_properties().size file_client.append_data(log_json + "\n", offset=current_length, length=len(log_json + "\n")) file_client.flush_data(current_length + len(log_json + "\n"))
内容的提问来源于stack exchange,提问作者Bee_Riii
相关产品推荐
相关产品推荐

