You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何通过单次API调用获取多个AWS Glue作业定义的最新运行元数据

解决方案:并发调用提升多Glue作业运行元数据获取效率

AWS Glue没有提供单次API调用即可获取多个作业最新运行元数据的原生接口,get_job_runs只能针对单个作业查询。不过可以通过并发调用API的方式,大幅提升批量获取的效率,替代逐个同步调用的低效方式。

实现思路

使用Python的concurrent.futures.ThreadPoolExecutor创建线程池,同时发起多个get_job_runs请求,每个请求对应一个目标作业,最后统一收集每个作业的最新运行记录。这种方式能把多个请求的等待时间重叠,总耗时接近单个请求的响应时间。

Boto3代码示例

import boto3
from concurrent.futures import ThreadPoolExecutor, as_completed

def get_latest_job_run(job_name):
    """获取单个Glue作业的最新运行元数据"""
    glue_client = boto3.client('glue')
    try:
        response = glue_client.get_job_runs(JobName=job_name, MaxResults=1)
        # 返回最新的运行记录(如果存在),否则返回None
        return {
            'job_name': job_name,
            'latest_run': response.get('JobRuns', [None])[0]
        }
    except Exception as e:
        return {
            'job_name': job_name,
            'error': str(e)
        }

def get_latest_runs_for_multiple_jobs(job_names, max_workers=5):
    """批量获取多个Glue作业的最新运行元数据"""
    results = []
    # 创建线程池,max_workers控制并发数(避免超过AWS API限额)
    with ThreadPoolExecutor(max_workers=max_workers) as executor:
        # 提交所有作业的查询任务
        future_to_job = {executor.submit(get_latest_job_run, job): job for job in job_names}
        # 逐个处理完成的任务
        for future in as_completed(future_to_job):
            results.append(future.result())
    return results

# 示例使用
if __name__ == '__main__':
    target_jobs = ['job-name-1', 'job-name-2', 'job-name-3']
    latest_runs = get_latest_runs_for_multiple_jobs(target_jobs, max_workers=5)
    
    # 打印结果
    for item in latest_runs:
        if 'error' in item:
            print(f"作业 {item['job_name']} 查询失败: {item['error']}")
        else:
            run = item['latest_run']
            if run:
                print(f"作业 {item['job_name']} 最新运行:")
                print(f"  运行ID: {run['Id']}")
                print(f"  状态: {run['JobRunState']}")
                print(f"  开始时间: {run['StartedOn']}")
            else:
                print(f"作业 {item['job_name']} 无运行记录")

注意事项

  • 并发数控制:max_workers不要设置过高,避免触发AWS Glue API的请求限额(默认限额可参考AWS官方文档,若需要可申请提升)。
  • 异常处理:代码中包含了基础异常捕获,可根据实际需求扩展(比如单独处理EntityNotFoundException即作业不存在的情况)。
  • 结果过滤:get_job_runs的MaxResults=1参数确保只返回最新的一条运行记录,减少数据传输量和处理成本。

内容的提问来源于stack exchange,提问作者Mohamed Ali

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.13 23:43:19