You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何提升Python中批量调用calculate函数的运行速度?

提速思路与实现

原代码运行缓慢的核心原因是单循环逐点处理+重复创建对象/解析参数,以下是针对性的提速方案:

1. 向量化计算(最优方案,效率提升最明显)

pvlib的底层太阳位置计算函数支持批量输入经纬度,无需逐个循环处理。直接生成经纬度网格一次性计算,彻底消除循环开销。

修改后的代码:

import numpy as np
import pvlib.solarposition
import pandas as pd

longitude_extent = (50, 60)
latitude_extent = (110, 120)

# 生成100×100的经纬度网格
longitudes = np.linspace(*longitude_extent, 100)
latitudes = np.linspace(*latitude_extent, 100)
lon_grid, lat_grid = np.meshgrid(longitudes, latitudes, indexing='ij')

# 提前转换时间为DatetimeIndex,避免重复解析字符串
dt = pd.DatetimeIndex(["2022-07-01 00:00:00"])

# 批量计算所有点的太阳天顶角
solar_pos = pvlib.solarposition.get_solarposition(
    time=dt,
    latitude=lat_grid.flatten(),
    longitude=lon_grid.flatten(),
    altitude=0,  # 与原代码Location的海拔参数一致
    tz='utc'
)

# 将结果重塑为100×100的数组
result = solar_pos['zenith'].values.reshape(100, 100)

说明:跳过Location实例化的冗余步骤,利用numpy向量化运算,速度可提升几十倍至上百倍。

2. 并行计算(适合无法向量化的场景)

如果必须保留逐点处理逻辑,可利用CPU多核并行计算10000个坐标点:

多线程实现(ThreadPoolExecutor)

import numpy as np
from concurrent.futures import ThreadPoolExecutor
from pvlib.location import Location

longitude_extent = (50, 60)
latitude_extent = (110, 120)
longitudes = np.linspace(*longitude_extent, 100)
latitudes = np.linspace(*latitude_extent, 100)
date_time = "2022-07-01 00:00:00"

def calculate(lat, lon):
    site = Location(lat, lon, 'utc', 0)
    return site.get_solarposition(date_time)['zenith']

# 生成所有坐标对
coords = [(lat, lon) for lon in longitudes for lat in latitudes]

# 多线程并行计算(线程数建议设为CPU核心数的1-2倍)
with ThreadPoolExecutor(max_workers=8) as executor:
    results = list(executor.map(lambda x: calculate(*x), coords))

# 重塑为100×100数组
result = np.array(results).reshape(100, 100)

说明:若计算为CPU密集型,可改用ProcessPoolExecutor规避Python GIL限制。

3. 减少重复开销的轻量优化

针对原代码的冗余操作做精简,无需大改结构即可提速:

import numpy as np
import pvlib.solarposition
import pandas as pd

longitude_extent = (50, 60)
latitude_extent = (110, 120)
longitudes = np.linspace(*longitude_extent, 100)
latitudes = np.linspace(*latitude_extent, 100)
# 提前解析时间,避免循环内重复转换
dt = pd.DatetimeIndex(["2022-07-01 00:00:00"])

def calculate(lat, lon):
    # 直接调用底层函数,跳过Location实例化
    solar_pos = pvlib.solarposition.get_solarposition(
        time=dt, latitude=lat, longitude=lon, altitude=0, tz='utc'
    )
    return solar_pos['zenith'].iloc[0]

result = np.zeros((100, 100))
for ix, lon in enumerate(longitudes):
    for jx, lat in enumerate(latitudes):
        result[ix, jx] = calculate(lat, lon)

说明:减少了Location实例化和时间字符串重复解析的开销,比原代码快2-5倍。

4. 其他细节优化

  • 删除未使用的导入:原代码导入的pandas(除时间处理外)、matplotlib.pyplot可直接删除,减少启动开销。
  • 用numpy数组操作替代纯Python循环:尽量避免嵌套循环中的Python原生运算,改用numpy广播机制。

内容的提问来源于stack exchange,提问作者mark

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.20 18:40:30