如何优化Python代码提升批量时序数据点的下载速度
时序数据下载效率优化方案
核心瓶颈说明
你当前的写法是一次性将187个ID全部传入单接口请求,这类长跨度时序数据拉取场景下,单请求体量过大极易触发服务端限流、单链路传输超时、服务端任务排队,是耗时过长的核心原因。
最高性价比优化方案:多线程分批次请求
IO密集型的网络请求场景,用多线程拆分请求即可获得数倍到数十倍的提速,实现成本极低:
import pandas as pd from concurrent.futures import ThreadPoolExecutor # 读取ID列表 df = pd.ExcelFile('ids.xlsx').parse('Sheet1') id_list = df['external_ids'].to_list() # 拆分为小批次,单批ID数量可根据接口限流情况调整,推荐初始设为5-20 batch_size = 10 id_batches = [id_list[i:i+batch_size] for i in range(0, len(id_list), batch_size)] # 单批次下载逻辑,可直接复用你原有请求参数 def download_batch(batch_ids): return client.datapoints.retrieve_dataframe( external_id=batch_ids, start=0, end="now" ) # 多线程并行下载,线程数推荐初始设为3-8,避免并发过高触发限流 all_batch_df = [] with ThreadPoolExecutor(max_workers=5) as executor: for batch_result in executor.map(download_batch, id_batches): all_batch_df.append(batch_result) # 合并所有批次结果,和你原有逻辑返回的dfdps格式完全一致 dfdps = pd.concat(all_batch_df, axis=1)
可选增强优化
- 加失败重试逻辑,避免偶发网络波动或限流导致的整体任务失败,示例如下(需要先安装
tenacity库:pip install tenacity):
from tenacity import retry, stop_after_attempt, wait_exponential # 最多重试3次,重试间隔按2/4/8秒递增 @retry(stop=stop_after_attempt(3), wait=wait_exponential(multiplier=1, min=2, max=10)) def download_batch(batch_ids): return client.datapoints.retrieve_dataframe( external_id=batch_ids, start=0, end="now" )
- 若单ID的时间跨度极大,可进一步将时间范围拆分为多个时间段,和ID批次组合拆分请求,进一步降低单请求体量。
内容的提问来源于stack exchange,提问作者ZZZSharePoint
相关产品推荐
相关产品推荐

