为何复杂Polars表达式单线程执行比多线程更快?
Polars >=0.17.13版本单线程性能反超多线程的特殊现象
令人意外的是,当仅限制使用1个线程时,Polars >= 0.17.13的性能表现反而更优;部分低版本Polars则不存在这种差异。
测试代码
# 适配polars>=0.17.3版本 import polars as pl import time import os def compute(initial_value: pl.Expr) -> pl.Expr: """通过循环计算消耗CPU资源,模拟黄金比例推导过程""" result = initial_value for i in range(100): result = (1 + result) ** 0.5 return result # 限制Polars表达式并行执行(线程数越多反而越慢) os.environ["POLARS_MAX_THREADS"] = "1" # 创建包含1000万条数据的DataFrame df = pl.DataFrame(pl.repeat(0.0, 10_000_000, eager=True).alias("value")) # 执行计算并统计耗时 start_time = time.perf_counter() df_b = ( df.lazy().with_columns(compute(pl.col("value")).alias("result")).collect() ) delta_time = time.perf_counter() - start_time print(f"线程池大小: {pl.threadpool_size()}. 耗时: {delta_time:.3f} s.")
运行结果对比
以下是分别设置与未设置os.environ["POLARS_MAX_THREADS"] = "1"时的运行结果:
线程池大小: 1. 耗时: 2.102 s. 线程池大小: 16. 耗时: 3.634 s.
内容的提问来源于stack exchange,提问作者Stas Stepanov
相关产品推荐
相关产品推荐

