如何让Polars搭配calamine引擎更优雅地处理崩溃?
如何捕获Polars + Calamine引擎触发的Rust Panic并优雅处理Excel读取失败
问题背景
需要批量验证上万份Excel文件,使用Polars的read_excel方法并指定engine="calamine"时,部分存在格式问题的Excel会触发Rust层面的index out of bounds panic,抛出pyo3_runtime.PanicException。由于该异常属于Python的BaseException子类而非普通Exception,常规的try-except Exception无法捕获,导致脚本直接崩溃,无法记录失败的文件名。已知这类异常文件无法避免,需要优雅处理并继续执行后续文件的验证。
最小复现代码:
from pathlib import Path import polars as pl root = Path("/path/to/globdir") def try_to_open(): for file in root.rglob("*/*_fileID.xlsx"): print(f"\r{file.name}", end='') try: df = pl.read_excel(file, engine="calamine", infer_schema_length=0) except Exception as e: print(f"{file.name}: {e}", flush=True) def main(): try_to_open() if __name__ == "__main__": main()
触发的错误信息:
thread '<unnamed>' panicked at /root/.cargo/registry/src/index.crates.io-6f17d22bba15001f/calamine-0.26.1/src/xlsx/cells_reader.rs:347:39: index out of bounds: the len is 2585 but the index is 2585 note: run with `RUST_BACKTRACE=1` environment variable to display a backtrace Traceback (most recent call last): File "/path/to/script.py", line 18, in <module> main() File "/path/to/script.py", line 15, in main try_to_open() File "/path/to/script.py", line 10, in try_to_open df = pl.read_excel(file, engine="calamine", infer_schema_length=0) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ ... pyo3_runtime.PanicException: index out of bounds: the len is 2585 but the index is 2585
解决方案
方案1:扩展异常捕获范围,直接捕获PanicException或BaseException
pyo3_runtime.PanicException属于BaseException子类,因此可以通过以下两种方式捕获:
方式A:明确捕获PanicException
需先导入pyo3_runtime模块(可通过pip install pyo3间接获取依赖):
from pathlib import Path import polars as pl from pyo3_runtime import PanicException root = Path("/path/to/globdir") def try_to_open(): for file in root.rglob("*/*_fileID.xlsx"): print(f"\r{file.name}", end='') try: df = pl.read_excel(file, engine="calamine", infer_schema_length=0) except Exception as e: print(f"\n{file.name}: {e}", flush=True) except PanicException as pe: print(f"\n{file.name}: Rust panic触发 - {pe}", flush=True) def main(): try_to_open() if __name__ == "__main__": main()
方式B:捕获所有BaseException(需排除中断信号)
如果无法导入PanicException,可以捕获BaseException,但要保留KeyboardInterrupt的处理,避免无法通过Ctrl+C终止程序:
from pathlib import Path import polars as pl root = Path("/path/to/globdir") def try_to_open(): for file in root.rglob("*/*_fileID.xlsx"): print(f"\r{file.name}", end='') try: df = pl.read_excel(file, engine="calamine", infer_schema_length=0) except KeyboardInterrupt: # 保留Ctrl+C的终止能力 raise except BaseException as e: print(f"\n{file.name}: 处理失败 - {e}", flush=True) def main(): try_to_open() if __name__ == "__main__": main()
方案2:用子进程隔离单个文件的处理
如果Rust panic可能导致进程状态不稳定,推荐用子进程单独处理每个Excel文件,单个子进程的崩溃不会影响主进程:
步骤1:编写子进程脚本process_excel.py
from pathlib import Path import polars as pl import sys def main(): if len(sys.argv) != 2: sys.exit(1) file_path = Path(sys.argv[1]) try: # 这里可以添加你的验证逻辑 df = pl.read_excel(file_path, engine="calamine", infer_schema_length=0) sys.exit(0) except Exception as e: print(f"普通错误: {e}", file=sys.stderr) sys.exit(1) except BaseException as be: print(f"Rust panic: {be}", file=sys.stderr) sys.exit(2) if __name__ == "__main__": main()
步骤2:主脚本调用子进程
from pathlib import Path import subprocess import sys root = Path("/path/to/globdir") def try_to_open(): # 记录失败文件的日志 with open("failed_excel_log.txt", "a") as log_file: for file in root.rglob("*/*_fileID.xlsx"): print(f"\r正在处理: {file.name}", end='', flush=True) # 调用子进程处理单个文件 result = subprocess.run( [sys.executable, "process_excel.py", str(file)], capture_output=True, text=True ) if result.returncode != 0: error_msg = result.stderr.strip() if result.stderr else f"未知错误,退出码: {result.returncode}" log_entry = f"{str(file)}: {error_msg}\n" print(f"\n处理失败: {file.name}", flush=True) log_file.write(log_entry) def main(): try_to_open() if __name__ == "__main__": main()
方案3:双引擎 fallback(优先Calamine,失败则用xlsx2csv)
已知xlsx2csv引擎不会触发此类panic,可以在Calamine失败时自动切换引擎重试,兼顾性能和兼容性:
from pathlib import Path import polars as pl from pyo3_runtime import PanicException root = Path("/path/to/globdir") def try_to_open(): for file in root.rglob("*/*_fileID.xlsx"): print(f"\r{file.name}", end='') success = False # 先尝试Calamine引擎 try: df = pl.read_excel(file, engine="calamine", infer_schema_length=0) # 执行验证逻辑 success = True except (Exception, PanicException) as e: # 失败则切换到xlsx2csv try: df = pl.read_excel(file, engine="xlsx2csv", infer_schema_length=0) # 执行验证逻辑 success = True print(f"\n{file.name}: 切换到xlsx2csv引擎处理成功", flush=True) except Exception as ee: print(f"\n{file.name}: 双引擎均处理失败 - Calamine错误: {e}, xlsx2csv错误: {ee}", flush=True) # 记录失败文件 if not success: with open("failed_files.txt", "a") as f: f.write(f"{str(file)}\n") def main(): try_to_open() if __name__ == "__main__": main()
注意事项
- 捕获
BaseException时务必排除KeyboardInterrupt,否则无法通过常规方式终止脚本。 - 子进程方案会带来一定的性能开销,但能彻底隔离panic对主进程的影响,适合稳定性要求高的场景。
- 如果选择双引擎fallback,需要确保环境中已安装
xlsx2csv(可通过pip install xlsx2csv安装)。
内容的提问来源于stack exchange,提问作者skytwosea
相关产品推荐
相关产品推荐

