You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python定时上传AWS S3遇网络异常的优雅处理方案咨询

老兄,你遇到的这个网络异常导致定时任务直接挂掉的问题,我之前也碰过好几次,用异常捕获+智能重试机制就能完美解决——既不会让整个程序崩掉,还能在网络恢复后自动续传。下面给你一步步拆解可行的实现方式:

1. 先把异常兜住,别让程序直接崩溃

首先,得把你job函数里的S3上传逻辑用try-except包裹起来,专门捕获网络相关的异常(包括你提到的socket.gaierror)。这样哪怕某次上传失败,也不会连累整个调度器进程直接终止。

修改后的job函数大概是这样:

import socket
import boto3
from botocore.exceptions import ClientError, ConnectionError

def job():
    serialNo = getserial()
    # 假设这里是你的文件准备逻辑
    file_path = f"device_{serialNo}_data.txt"
    
    try:
        # 初始化S3客户端
        s3 = boto3.client('s3')
        # 执行上传
        s3.upload_file(file_path, 'your-target-bucket', f'device-data/{file_path}')
        print(f"✅ 文件 {file_path} 上传成功")
    except socket.gaierror as e:
        print(f"⚠️ 网络域名解析失败,稍后重试: {str(e)}")
        # 触发重试逻辑
        retry_upload(file_path)
    except (ConnectionError, ClientError) as e:
        # 兜底捕获其他S3连接/客户端异常
        print(f"⚠️ S3上传失败: {str(e)}")
        retry_upload(file_path)
    except Exception as e:
        # 最后兜底所有未预料的异常,绝对不让程序崩掉
        print(f"❌ 未知错误: {str(e)}")

2. 用指数退避重试,别死磕浪费资源

直接反复重试太粗暴,万一网络一直不好,会白白消耗CPU和带宽。推荐用指数退避的方式:每次重试的间隔时间翻倍,直到成功或者达到最大重试次数。

方式一:自己手动实现(不用额外装库)

import time

def retry_upload(file_path, max_retries=5):
    retry_count = 0
    while retry_count < max_retries:
        try:
            s3 = boto3.client('s3')
            s3.upload_file(file_path, 'your-target-bucket', f'device-data/{file_path}')
            print(f"🔄 重试上传 {file_path} 成功")
            return
        except (socket.gaierror, ConnectionError, ClientError) as e:
            retry_count += 1
            # 指数退避:2^1, 2^2...最多到30秒(避免间隔太长)
            wait_time = min(2 ** retry_count, 30)
            print(f"🔄 第 {retry_count} 次重试,等待 {wait_time} 秒... 错误: {str(e)}")
            time.sleep(wait_time)
    # 达到最大重试次数,记录下来后续手动处理
    print(f"❌ 重试次数耗尽,文件 {file_path} 上传失败,请稍后检查")

方式二:用tenacity库简化(更优雅)

先装库:pip install tenacity

然后用装饰器快速实现重试逻辑:

from tenacity import retry, stop_after_attempt, wait_exponential, retry_if_exception_type

# 定义带重试的上传函数
@retry(
    stop=stop_after_attempt(5),  # 最多重试5次
    wait=wait_exponential(multiplier=1, min=2, max=30),  # 指数退避,最小2秒,最大30秒
    retry=retry_if_exception_type((socket.gaierror, ConnectionError, ClientError))
)
def upload_to_s3(file_path):
    s3 = boto3.client('s3')
    s3.upload_file(file_path, 'your-target-bucket', f'device-data/{file_path}')

# 改造后的job函数
def job():
    serialNo = getserial()
    file_path = f"device_{serialNo}_data.txt"
    try:
        upload_to_s3(file_path)
        print(f"✅ 文件 {file_path} 上传成功")
    except Exception as e:
        print(f"❌ 多次重试后仍失败: {str(e)}")

3. 确保调度器本身不崩

如果你用的是APScheduler这类调度库,还要给调度器加一层异常防护,避免单个job的异常把整个调度器带崩。比如APScheduler的启动逻辑可以这么写:

from apscheduler.schedulers.background import BackgroundScheduler
import time

def start_scheduler():
    scheduler = BackgroundScheduler()
    # 假设你是每天凌晨2点执行任务
    scheduler.add_job(job, 'cron', hour=2)
    try:
        scheduler.start()
        print("🚀 调度器启动成功")
        # 保持主进程运行
        while True:
            time.sleep(1)
    except (KeyboardInterrupt, SystemExit):
        scheduler.shutdown()
        print("🛑 调度器已关闭")
    except Exception as e:
        print(f"⚠️ 调度器运行异常: {str(e)}")
        # 这里可以加告警逻辑,比如发邮件/企业微信通知

可选:加日志记录,方便排查问题

把print换成Python的logging模块,把所有操作和异常都记录到日志文件里,后续排查问题更方便:

import logging

logging.basicConfig(
    level=logging.INFO,
    format='%(asctime)s - %(levelname)s - %(message)s',
    filename='s3_upload.log'
)

# 替换所有print为logging
logging.info(f"✅ 文件 {file_path} 上传成功")
logging.warning(f"🔄 第 {retry_count} 次重试,等待 {wait_time} 秒... 错误: {str(e)}")
logging.error(f"❌ 多次重试后仍失败: {str(e)}")

这样处理后,下次再遇到网络波动,程序会自动进入重试流程,等网络恢复后就会自动完成上传,完全不用手动重启程序。

内容的提问来源于stack exchange,提问作者Marimuthu

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 09:53:06