You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何优化Python代码提升批量时序数据点的下载速度

时序数据下载效率优化方案

核心瓶颈说明

你当前的写法是一次性将187个ID全部传入单接口请求,这类长跨度时序数据拉取场景下,单请求体量过大极易触发服务端限流、单链路传输超时、服务端任务排队,是耗时过长的核心原因。

最高性价比优化方案:多线程分批次请求

IO密集型的网络请求场景,用多线程拆分请求即可获得数倍到数十倍的提速,实现成本极低:

import pandas as pd
from concurrent.futures import ThreadPoolExecutor

# 读取ID列表
df = pd.ExcelFile('ids.xlsx').parse('Sheet1')
id_list = df['external_ids'].to_list()

# 拆分为小批次,单批ID数量可根据接口限流情况调整,推荐初始设为5-20
batch_size = 10
id_batches = [id_list[i:i+batch_size] for i in range(0, len(id_list), batch_size)]

# 单批次下载逻辑,可直接复用你原有请求参数
def download_batch(batch_ids):
    return client.datapoints.retrieve_dataframe(
        external_id=batch_ids,
        start=0,
        end="now"
    )

# 多线程并行下载,线程数推荐初始设为3-8,避免并发过高触发限流
all_batch_df = []
with ThreadPoolExecutor(max_workers=5) as executor:
    for batch_result in executor.map(download_batch, id_batches):
        all_batch_df.append(batch_result)

# 合并所有批次结果,和你原有逻辑返回的dfdps格式完全一致
dfdps = pd.concat(all_batch_df, axis=1)

可选增强优化

  • 加失败重试逻辑,避免偶发网络波动或限流导致的整体任务失败,示例如下(需要先安装tenacity库:pip install tenacity):
from tenacity import retry, stop_after_attempt, wait_exponential

# 最多重试3次,重试间隔按2/4/8秒递增
@retry(stop=stop_after_attempt(3), wait=wait_exponential(multiplier=1, min=2, max=10))
def download_batch(batch_ids):
    return client.datapoints.retrieve_dataframe(
        external_id=batch_ids,
        start=0,
        end="now"
    )
  • 若单ID的时间跨度极大,可进一步将时间范围拆分为多个时间段,和ID批次组合拆分请求,进一步降低单请求体量。

内容的提问来源于stack exchange,提问作者ZZZSharePoint

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.02 02:54:04