Python执行期间如何捕获模块日志输出并处理相关错误
问题:捕获cfgrib跳过变量的日志信息且不中断数据集加载
我在脚本中配置了日志记录器,能把__main__及所用模块的日志输出到标准输出和日志文件。执行xarray.open_dataset(file, engine="cfgrib")时,部分场景下cfgrib模块会触发DatasetBuildError,但这个错误被模块内部处理了,程序能继续运行,但日志里会出现类似这样的输出:
2023-02-18 10:02:06,731 cfgrib.dataset ERROR skipping variable: paramId==228029 shortName='i10fg'
我需要获取这类错误里的shortName信息来完成完整的错误处理,但直接用try-except抓不到这个异常。添加errors='raise'参数虽然能捕获异常,但会导致xr.open_dataset()调用失败,无法获取已经解析好的变量——这个调用耗时极长,重复调用成本太高。需要一种不中断调用、保留返回结果的前提下,实时判断是否有这类错误日志产生的方法,以便仅在有错误时补充调用获取剩余变量。
解决方案
1. 自定义日志处理器捕获特定日志
编写一个自定义日志Handler,专门监听cfgrib.dataset的ERROR级日志,提取其中的shortName信息。这种方式不会干扰原有日志输出,还能实时捕获目标信息。
import logging from collections import deque import xarray as xr # 存储捕获到的被跳过变量的shortName skipped_short_names = deque() class CfgribErrorCaptureHandler(logging.Handler): def emit(self, record): # 过滤目标日志:仅处理cfgrib.dataset的ERROR日志,且内容包含"skipping variable" if (record.name == "cfgrib.dataset" and record.levelno == logging.ERROR and "skipping variable" in record.getMessage()): msg = record.getMessage() # 从日志内容中提取shortName start_idx = msg.find("shortName='") + len("shortName='") end_idx = msg.find("'", start_idx) if start_idx != -1 and end_idx != -1: short_name = msg[start_idx:end_idx] skipped_short_names.append(short_name) # 给cfgrib.dataset日志器添加自定义处理器 cfgrib_logger = logging.getLogger("cfgrib.dataset") capture_handler = CfgribErrorCaptureHandler() cfgrib_logger.addHandler(capture_handler) # 执行数据集加载,不中断流程 ds = xr.open_dataset(file, engine="cfgrib") # 加载完成后检查是否有被跳过的变量 if skipped_short_names: print(f"检测到被跳过的变量:{list(skipped_short_names)}") # 补充加载这些变量 for short_name in skipped_short_names: try: # 用filter_by_keys仅加载目标变量,减少耗时 var_ds = xr.open_dataset(file, engine="cfgrib", filter_by_keys={"shortName": short_name}) ds = ds.merge(var_ds) except Exception as e: print(f"加载变量{short_name}失败:{str(e)}") # 清理自定义处理器(可选) cfgrib_logger.removeHandler(capture_handler)
2. 临时重定向日志流解析内容
如果不想修改全局日志配置,可以临时将cfgrib.dataset的日志输出重定向到内存流,加载完成后再解析日志内容提取信息。
import io import logging import xarray as xr from contextlib import contextmanager @contextmanager def capture_cfgrib_errors(): log_stream = io.StringIO() # 创建临时日志处理器 temp_handler = logging.StreamHandler(log_stream) temp_handler.setLevel(logging.ERROR) # 保存cfgrib日志器原有处理器 cfgrib_logger = logging.getLogger("cfgrib.dataset") original_handlers = cfgrib_logger.handlers.copy() cfgrib_logger.handlers = [temp_handler] try: yield log_stream finally: # 恢复原有日志处理器 cfgrib_logger.handlers = original_handlers # 使用上下文管理器捕获日志 with capture_cfgrib_errors() as log_stream: ds = xr.open_dataset(file, engine="cfgrib") # 解析日志内容 skipped_short_names = [] log_content = log_stream.getvalue() for line in log_content.splitlines(): if "ERROR skipping variable" in line: msg = line.split("ERROR ")[1] start_idx = msg.find("shortName='") + len("shortName='") end_idx = msg.find("'", start_idx) if start_idx != -1 and end_idx != -1: skipped_short_names.append(msg[start_idx:end_idx]) # 处理被跳过的变量 if skipped_short_names: # 执行补充加载逻辑 pass
关键注意点
- 提取
shortName的逻辑需匹配实际日志格式,若后续cfgrib日志格式变更,需同步调整解析代码。 - 补充加载变量时,使用
filter_by_keys参数仅加载目标变量,避免重复全量读取数据集,大幅降低耗时。 - 自定义Handler方式适合长期监控日志,临时重定向方式更适合单次调用场景。
内容的提问来源于stack exchange,提问作者squarespiral
相关产品推荐
相关产品推荐

