You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python含200个ID的列表如何分批次调用并优化多线程执行效率?

解决方案

现有代码问题分析

  • 你当前的代码虽然初始化了最大20线程的线程池,但仅提交了1个任务,一次性请求全部200个ID,完全没有利用多线程并行能力,这是运行耗时过长的核心原因
  • ID列表存储逻辑冗余,用全局变量传递参数也容易引发并发场景下的异常
  • 缺少异常捕获逻辑,单批次请求失败会导致整个任务终止

实现分批次调用+多线程优化的完整代码

from concurrent.futures import ThreadPoolExecutor, as_completed
import pandas as pd
import numpy as np
import datetime as dt

# 读取ID列表
df = pd.ExcelFile('ids.xlsx').parse('Sheet1')
id_list = df['external_ids'].to_list()

# 分批函数,可调整batch_size为10或20
def split_batches(lst, batch_size=20):
    for i in range(0, len(lst), batch_size):
        yield lst[i:i+batch_size]

# 下载函数改为接收批次ID作为参数
def download(batch_ids):
    try:
        # client是你自己的SDK实例,这里保留你的原有逻辑
        dtest_df = client.datapoints.retrieve_dataframe(
            external_id = batch_ids, 
            start=0, 
            end="now",
            granularity='1m'
        )        
        dtest_df = dtest_df.rename(columns = {'index':'timestamp'})
        client.datapoints.insert_dataframe(
            dtest_df,
            external_id_headers = True,
            dropna = True
        )
        return f"批次处理完成,共{len(batch_ids)}个ID"
    except Exception as e:
        return f"批次处理失败,错误信息:{str(e)}"

if __name__ == "__main__":
    batches = list(split_batches(id_list, batch_size=20))
    # 线程数可根据API限流规则调整,避免并发过高被封禁
    max_workers = 5
    with ThreadPoolExecutor(max_workers=max_workers) as executor:
        futures = [executor.submit(download, batch) for batch in batches]
        for future in as_completed(futures):
            print(future.result())

额外优化建议

  • 你可以根据目标API的限流规则调整batch_size和max_workers参数,如果API允许更高并发可以调大max_workers,如果频繁触发限流就调小参数
  • 如果插入数据的操作耗时很高,可以考虑把查询和插入拆分为两个独立的线程池,实现生产消费模式进一步提速
  • 可以给每个批次加重试逻辑,应对偶发的网络波动或者API超时问题

内容的提问来源于stack exchange,提问作者ZZZSharePoint

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.01 04:18:03