You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

关于scipy.interpolate.RegularGridInterpolator的异常耗时问题

分块插值提升RegularGridInterpolator性能的原因分析

现象描述

使用scipy.interpolate.RegularGridInterpolator对大数组插值时,将数组拆分为多块分别插值再拼接结果的速度反而更快。例如对5维100万个点插值时,分块插值速度几乎是单次插值的两倍。

测试代码

from scipy.interpolate import RegularGridInterpolator
import numpy as np
import time

N = [10, 11, 12, 13, 14]
grid = [np.linspace(i, i + 1, n) for i, n in enumerate(N, 0)]
values = np.arange(np.prod(N)).reshape(N)
rgi = RegularGridInterpolator(grid, values)

for N in [10000, 100000, 1000000]:
    xp = np.array([np.random.random(N) + i for i in range(5)]).T

    t = time.time()
    r1 = rgi(xp)
    t1 = time.time() - t
    t = time.time()
    r2 = np.concatenate([rgi(xp_) for xp_ in np.array_split(xp, 100)])
    t2 = time.time() - t
    print(f'{N}: {t1:.3f}, 100x{N/100}: {t2:.3f}')

测试输出

10000: 0.009, 100x100.0: 0.051
100000: 0.087, 100x1000.0: 0.067
1000000: 1.094, 100x10000.0: 0.594

环境版本

Python 3.13.9
numpy   2.3.4
pip     25.3
scipy   1.16.3

原因解析

这个现象的核心是CPU缓存利用率的差异:

  • 单次处理大数组时,插值过程中产生的中间数据量远超CPU缓存容量,导致频繁的缓存失效(cache miss),大量时间消耗在内存与缓存的数据交换上,拖慢整体速度。
  • 分块处理时,每块数据规模小,能完全适配CPU缓存,缓存命中率大幅提升,计算效率显著提高。虽然分块会带来额外的函数调用和数组拼接开销,但当数据量足够大时,缓存优化的收益远超过这些额外成本。

此外,RegularGridInterpolator的内部实现对超大规模输入的内存分配和处理逻辑未做针对性优化,分块后每一次调用的内存压力更小,处理流程更高效。

总结

当插值点数超过一定阈值时,分块处理是提升RegularGridInterpolator性能的有效手段。最优分块大小可根据实际数据维度和硬件环境调整,不一定局限于100块。

内容的提问来源于stack exchange,提问作者Mads M Pedersen

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.12 01:20:11